Datasets:
Make release totals registry-driven (#29)
Browse files- fix: make release totals registry-driven, not "whatever is on main" (4f879612263127710c44a0c376b697088f2ca146)
- docs: how to push and open PRs on the Hub (10da95c3d913e08491106f1b6628b59e15105738)
- feat: audit source licenses and table footnotes (4fbe51379b807370e4268e0a7b2809ffc5ac43da)
- feat: verify the license-family table is a complete partition (d0b620e22be4c888ae3851602ba012961eb61245)
Co-authored-by: Paweł Puzio <ppuzio@users.noreply.huggingface.co>
- AGENTS.md +53 -7
- CHANGELOG.md +12 -2
- README.md +10 -1
- src/make_docs.py +61 -14
- src/sources.py +48 -0
- src/test_verify_paper_claims.py +190 -0
- src/verify_paper_claims.py +637 -0
AGENTS.md
CHANGED
|
@@ -13,8 +13,9 @@ committed docs on its own:
|
|
| 13 |
|
| 14 |
- It iterates `SOURCES` in `src/sources.py`; any source in `data/` but missing
|
| 15 |
from `SOURCES` is silently dropped from the table and totals.
|
| 16 |
-
- It stamps
|
| 17 |
-
constant
|
|
|
|
| 18 |
- It cannot reproduce hand-written README narrative (audit notes, policy notes,
|
| 19 |
the phrase-frequency section).
|
| 20 |
|
|
@@ -44,6 +45,11 @@ answer was yes or no.
|
|
| 44 |
`speakleash_key`/`file_key`. If the datasheet is hand-authored (rich provenance
|
| 45 |
or a fixed add date), add `"custom_datasheet": True` so make_docs leaves it
|
| 46 |
alone.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 47 |
4. Add a contract test: `src/test_<key>_contract.py` (canonical schema, non-empty
|
| 48 |
text, positive token counts, uniform source/license, stats-file consistency).
|
| 49 |
Run `python3 -m pytest src/` before committing.
|
|
@@ -142,6 +148,47 @@ Two rules about all of the above:
|
|
| 142 |
4. Prepend the release notes to `CHANGELOG.md` under a new `## vX.Y.Z (date)`
|
| 143 |
heading. History is append-only — never rewrite earlier releases.
|
| 144 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 145 |
## Do not commit
|
| 146 |
|
| 147 |
`*.log` (fetch/build logs), `.DS_Store`, `src/__pycache__/`. Only `logs/` is in
|
|
@@ -150,8 +197,7 @@ before staging. `*.parquet` is LFS-tracked; commit the pointer, not the blob.
|
|
| 150 |
|
| 151 |
## Known drift / cleanup opportunities
|
| 152 |
|
| 153 |
-
- `
|
| 154 |
-
`
|
| 155 |
-
|
| 156 |
-
|
| 157 |
-
add date (or `custom_datasheet`) before making make_docs regenerate them.
|
|
|
|
| 13 |
|
| 14 |
- It iterates `SOURCES` in `src/sources.py`; any source in `data/` but missing
|
| 15 |
from `SOURCES` is silently dropped from the table and totals.
|
| 16 |
+
- It stamps a regenerated datasheet's "Added" with the source's `added` field,
|
| 17 |
+
falling back to the global `ADDED` constant when the entry has none — so a new
|
| 18 |
+
entry without `added` still gets a wrong date.
|
| 19 |
- It cannot reproduce hand-written README narrative (audit notes, policy notes,
|
| 20 |
the phrase-frequency section).
|
| 21 |
|
|
|
|
| 45 |
`speakleash_key`/`file_key`. If the datasheet is hand-authored (rich provenance
|
| 46 |
or a fixed add date), add `"custom_datasheet": True` so make_docs leaves it
|
| 47 |
alone.
|
| 48 |
+
|
| 49 |
+
Also set `added` (the real add date) and `release`. `release` is what admits a
|
| 50 |
+
source to a release's totals: `None` means "on main, not in any release yet",
|
| 51 |
+
which is the correct value for a new source until the release that ships it is
|
| 52 |
+
cut. Nothing is counted just because it has a `stats.json`.
|
| 53 |
4. Add a contract test: `src/test_<key>_contract.py` (canonical schema, non-empty
|
| 54 |
text, positive token counts, uniform source/license, stats-file consistency).
|
| 55 |
Run `python3 -m pytest src/` before committing.
|
|
|
|
| 148 |
4. Prepend the release notes to `CHANGELOG.md` under a new `## vX.Y.Z (date)`
|
| 149 |
heading. History is append-only — never rewrite earlier releases.
|
| 150 |
|
| 151 |
+
## Push to Hugging Face
|
| 152 |
+
|
| 153 |
+
`origin` is the Hub itself
|
| 154 |
+
(`https://huggingface.co/datasets/SlayerLab/polish-dynaword`) — there is no GitHub
|
| 155 |
+
remote. `git push` authenticates through the git credential helper (macOS:
|
| 156 |
+
`osxkeychain`), which is separate from `hf auth login`: pushes can work fine while
|
| 157 |
+
`hf auth whoami` still reports "Not logged in". The Python API and `hf` CLI need
|
| 158 |
+
the login (or `HF_TOKEN`).
|
| 159 |
+
|
| 160 |
+
Hub pull requests use no forks and no named branches. A PR *is* the ref
|
| 161 |
+
`refs/pr/N` and the Hub assigns `N`, so you cannot open one by pushing — the ref
|
| 162 |
+
has to exist first. (Plain branches can be pushed, but a branch is not a PR and
|
| 163 |
+
nobody reviews it.)
|
| 164 |
+
|
| 165 |
+
Open a PR for work already committed locally:
|
| 166 |
+
|
| 167 |
+
```bash
|
| 168 |
+
# 1. create the empty PR (needs the write token)
|
| 169 |
+
python3 -c "
|
| 170 |
+
from huggingface_hub import HfApi
|
| 171 |
+
pr = HfApi().create_pull_request(
|
| 172 |
+
'SlayerLab/polish-dynaword', repo_type='dataset',
|
| 173 |
+
title='<title>', description='<what changed and why>')
|
| 174 |
+
print(pr.num, pr.url)"
|
| 175 |
+
|
| 176 |
+
# 2. push your local branch onto that ref
|
| 177 |
+
git push origin <local-branch>:refs/pr/<N>
|
| 178 |
+
```
|
| 179 |
+
|
| 180 |
+
Opened this way the PR starts in **draft** — publish it from the web UI.
|
| 181 |
+
|
| 182 |
+
Pick up an existing PR (42 here):
|
| 183 |
+
|
| 184 |
+
```bash
|
| 185 |
+
git fetch origin refs/pr/42:pr/42 && git checkout pr/42
|
| 186 |
+
git push origin pr/42:refs/pr/42
|
| 187 |
+
```
|
| 188 |
+
|
| 189 |
+
`repo_type='dataset'` is required on every `huggingface_hub` call — this is a
|
| 190 |
+
dataset repo, not a model. Push straight to `main` only when cutting a release.
|
| 191 |
+
|
| 192 |
## Do not commit
|
| 193 |
|
| 194 |
`*.log` (fetch/build logs), `.DS_Store`, `src/__pycache__/`. Only `logs/` is in
|
|
|
|
| 197 |
|
| 198 |
## Known drift / cleanup opportunities
|
| 199 |
|
| 200 |
+
- `europeana` is in `SOURCES` but has no `data/` shard anywhere, and
|
| 201 |
+
`parlamint_pl` has one only in some working trees — it is not committed.
|
| 202 |
+
make_docs skips both with `! no stats`. Intentional placeholders or stale —
|
| 203 |
+
confirm before relying on `build_dynaword.py --all`.
|
|
|
CHANGELOG.md
CHANGED
|
@@ -1,5 +1,14 @@
|
|
| 1 |
# Changelog
|
| 2 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 3 |
## v0.2.5 (2026-08-14)
|
| 4 |
|
| 5 |
- Added the community contribution `samorzad_gov_pl` by Dawid Majewski
|
|
@@ -43,8 +52,9 @@
|
|
| 43 |
- Expanded contributor credit for Arkadiusz Słota (`Maggio33`) for the WDS 10-5
|
| 44 |
expansion, PII-scrub and dedup work.
|
| 45 |
- Updated the complete-corpus release totals from per-source parquet stats:
|
| 46 |
-
4,246,429 documents and 9,597,908,067 cl100k-proxy tokens
|
| 47 |
-
separately-merged `biblioteka_nauki` source now present on
|
|
|
|
| 48 |
|
| 49 |
## v0.2.3 (2026-07-26)
|
| 50 |
|
|
|
|
| 1 |
# Changelog
|
| 2 |
|
| 3 |
+
## Unreleased (on main)
|
| 4 |
+
|
| 5 |
+
- `wiktionary_examples`, `sejm_api` and `parlamint_pl` are built, documented and
|
| 6 |
+
present on the main branch but have never been admitted to a stable release.
|
| 7 |
+
They are excluded from every release total until a release entry admits them.
|
| 8 |
+
The registry records this as `release: None` in `src/sources.py`; release
|
| 9 |
+
totals count only sources whose `release` is set and not newer than the
|
| 10 |
+
version being documented.
|
| 11 |
+
|
| 12 |
## v0.2.5 (2026-08-14)
|
| 13 |
|
| 14 |
- Added the community contribution `samorzad_gov_pl` by Dawid Majewski
|
|
|
|
| 52 |
- Expanded contributor credit for Arkadiusz Słota (`Maggio33`) for the WDS 10-5
|
| 53 |
expansion, PII-scrub and dedup work.
|
| 54 |
- Updated the complete-corpus release totals from per-source parquet stats:
|
| 55 |
+
4,246,429 documents and 9,597,908,067 cl100k-proxy tokens across 17 sources
|
| 56 |
+
(also reflects the separately-merged `biblioteka_nauki` source now present on
|
| 57 |
+
main).
|
| 58 |
|
| 59 |
## v0.2.3 (2026-07-26)
|
| 60 |
|
README.md
CHANGED
|
@@ -182,7 +182,16 @@ files is not listed as a data contribution.
|
|
| 182 |
| [wikinews](data/wikinews/wikinews.md) | Polish Wikinews | `CC-BY-2.5` | 24,386 | 12.1M |
|
| 183 |
| [global_voices](data/global_voices/global_voices.md) | Global Voices Polish | `CC-BY-3.0` | 2,040 | 3.7M |
|
| 184 |
| [nkjp1m](data/nkjp1m/nkjp1m.md) | The manually annotated 1-million word subcorpus of the National Corpus of Polish | `CC-BY` | 18,381 | 2.6M |
|
| 185 |
-
| **total** | | | **4,
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 186 |
|
| 187 |
## Method
|
| 188 |
Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
|
|
|
|
| 182 |
| [wikinews](data/wikinews/wikinews.md) | Polish Wikinews | `CC-BY-2.5` | 24,386 | 12.1M |
|
| 183 |
| [global_voices](data/global_voices/global_voices.md) | Global Voices Polish | `CC-BY-3.0` | 2,040 | 3.7M |
|
| 184 |
| [nkjp1m](data/nkjp1m/nkjp1m.md) | The manually annotated 1-million word subcorpus of the National Corpus of Polish | `CC-BY` | 18,381 | 2.6M |
|
| 185 |
+
| **total** | | | **4,319,200** | **9,639.1M** |
|
| 186 |
+
|
| 187 |
+
## Sources on main, not yet released
|
| 188 |
+
Built and documented, but **not** part of v0.2.5 and excluded from every total above. They join a release when a CHANGELOG entry admits them.
|
| 189 |
+
|
| 190 |
+
| source | description | license | documents | tokens |
|
| 191 |
+
|---|---|---|---:|---:|
|
| 192 |
+
| [parlamint_pl](data/parlamint_pl/parlamint_pl.md) | ParlaMint-PL (parliamentary debates, 2015-2022) | `CC-BY-4.0` | 686 | 95.4M |
|
| 193 |
+
| [sejm_api](data/sejm_api/sejm_api.md) | Sejm API parliamentary speeches (2023 onward) | `public-domain (official documents)` | 38,812 | 36.3M |
|
| 194 |
+
| [wiktionary_examples](data/wiktionary_examples/wiktionary_examples.md) | Polish Wiktionary usage examples | `CC-BY-SA-3.0` | 8,639 | 1.1M |
|
| 195 |
|
| 196 |
## Method
|
| 197 |
Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
|
src/make_docs.py
CHANGED
|
@@ -2,6 +2,9 @@
|
|
| 2 |
"""Generate Dynaword documentation: per-source datasheets + README + CHANGELOG + LICENSE.
|
| 3 |
|
| 4 |
Reads sources.py + data/<source>/<source>.stats.json (written by build_dynaword.py).
|
|
|
|
|
|
|
|
|
|
| 5 |
Implements the "Documented" principle (datasheets, Gebru et al. 2021) and the
|
| 6 |
aggregate README table (paper 2508.02271).
|
| 7 |
"""
|
|
@@ -49,12 +52,31 @@ from the next version (see retroactive-removal policy below).
|
|
| 49 |
"""
|
| 50 |
|
| 51 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 52 |
def load_stats(name):
|
| 53 |
f = ROOT / "data" / name / f"{name}.stats.json"
|
| 54 |
return json.loads(f.read_text()) if f.exists() else None
|
| 55 |
|
| 56 |
|
| 57 |
def datasheet(name, cfg, st):
|
|
|
|
| 58 |
license_rows = ""
|
| 59 |
licenses = st.get("licenses") or {}
|
| 60 |
if licenses:
|
|
@@ -84,7 +106,7 @@ def datasheet(name, cfg, st):
|
|
| 84 |
- **Language:** Polish (pl)
|
| 85 |
- **License:** `{cfg['license']}`
|
| 86 |
- **Created (range):** {cfg['created']}
|
| 87 |
-
- **Added:** {
|
| 88 |
|
| 89 |
## Licensing — traceable basis
|
| 90 |
{cfg['traceable']}
|
|
@@ -115,7 +137,8 @@ Llama-3 count is computed at release.
|
|
| 115 |
|
| 116 |
|
| 117 |
def main():
|
| 118 |
-
|
|
|
|
| 119 |
for name, cfg in SOURCES.items():
|
| 120 |
st = load_stats(name)
|
| 121 |
if not st:
|
|
@@ -123,16 +146,36 @@ def main():
|
|
| 123 |
# Sources with a hand-authored datasheet (rich provenance, per-source add
|
| 124 |
# date) opt out of template regeneration to avoid clobbering it.
|
| 125 |
if not cfg.get("custom_datasheet"):
|
| 126 |
-
(ROOT / "data" / name / f"{name}.md"
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 127 |
rows.append((name, cfg, st))
|
| 128 |
tot_doc += st["kept"]; tot_tok += st["tokens"]; tot_chr += st["chars"]
|
| 129 |
rows.sort(key=lambda r: -r[2]["tokens"])
|
|
|
|
| 130 |
|
| 131 |
tbl = "\n".join(
|
| 132 |
f"| [{n}](data/{n}/{n}.md) | {c['pretty']} | `{c['license']}` | "
|
| 133 |
f"{s['kept']:,} | {s['tokens']/1e6:,.1f}M |"
|
| 134 |
for n, c, s in rows)
|
| 135 |
excl = "\n".join(f"| `{k}` | {v} |" for k, v in EXCLUDED.items())
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 136 |
phrase_frequency = ""
|
| 137 |
phrase_path = ROOT / "artifacts" / "pattern_frequency_hf_snippet.md"
|
| 138 |
if phrase_path.exists():
|
|
@@ -143,7 +186,7 @@ def main():
|
|
| 143 |
"\n## Results\n\n"
|
| 144 |
"### Corpus phrase frequency (normalized by tokens)\n\n"
|
| 145 |
"Raw counts and token-normalized shares are regenerated from the "
|
| 146 |
-
"current
|
| 147 |
f"{snippet}"
|
| 148 |
)
|
| 149 |
|
|
@@ -206,11 +249,11 @@ family (Enevoldsen et al., [arXiv:2508.02271](https://arxiv.org/abs/2508.02271))
|
|
| 206 |
|
| 207 |
#### v0.2.5 — current stable
|
| 208 |
|
| 209 |
-
Added `samorzad_gov_pl`,
|
| 210 |
-
Polish local public institutions and one central platform
|
| 211 |
-
contribution preserves the publishing institution as `author`,
|
| 212 |
-
fetched article URLs in an attribution sidecar and the
|
| 213 |
-
basis.
|
| 214 |
|
| 215 |
#### v0.2.4 — previous stable
|
| 216 |
|
|
@@ -315,7 +358,7 @@ files is not listed as a data contribution.
|
|
| 315 |
|---|---|---|---:|---:|
|
| 316 |
{tbl}
|
| 317 |
| **total** | | | **{tot_doc:,}** | **{tot_tok/1e6:,.1f}M** |
|
| 318 |
-
|
| 319 |
## Method
|
| 320 |
Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
|
| 321 |
data. Gates are intentionally minimal (drop short docs, non-Polish, exact duplicates,
|
|
@@ -389,7 +432,7 @@ python3 src/make_docs.py
|
|
| 389 |
```
|
| 390 |
{phrase_frequency}
|
| 391 |
"""
|
| 392 |
-
(ROOT / "README.md"
|
| 393 |
|
| 394 |
# CHANGELOG is append-only release history. Documentation regeneration must
|
| 395 |
# never erase earlier releases or hand-reviewed legal/release notes.
|
|
@@ -403,15 +446,19 @@ python3 src/make_docs.py
|
|
| 403 |
f"{tot_doc:,} docs, {tot_tok:,} tokens (tiktoken cl100k proxy).\n\n"
|
| 404 |
)
|
| 405 |
changelog = changelog.replace("# Changelog\n", f"# Changelog\n\n{entry}", 1)
|
| 406 |
-
|
| 407 |
|
| 408 |
-
|
|
|
|
| 409 |
"Polish DynaWord is released under Creative Commons Attribution-ShareAlike 4.0\n"
|
| 410 |
"International (CC-BY-SA-4.0): https://creativecommons.org/licenses/by-sa/4.0/\n\n"
|
| 411 |
"Per-source upstream licenses and attribution are documented in each\n"
|
| 412 |
"data/<source>/<source>.md datasheet.\n")
|
| 413 |
|
| 414 |
-
print(f"docs written: README + CHANGELOG + LICENSE +
|
|
|
|
|
|
|
|
|
|
| 415 |
print(f"TOTAL {tot_doc:,} docs | {tot_tok/1e9:.2f}B tok | {tot_chr/1e9:.1f}B chars")
|
| 416 |
|
| 417 |
|
|
|
|
| 2 |
"""Generate Dynaword documentation: per-source datasheets + README + CHANGELOG + LICENSE.
|
| 3 |
|
| 4 |
Reads sources.py + data/<source>/<source>.stats.json (written by build_dynaword.py).
|
| 5 |
+
Only sources whose registry `release` is set and <= VERSION count toward release
|
| 6 |
+
totals; sources on main awaiting a release are listed separately and excluded, so
|
| 7 |
+
merging a new source cannot silently redefine a published total.
|
| 8 |
Implements the "Documented" principle (datasheets, Gebru et al. 2021) and the
|
| 9 |
aggregate README table (paper 2508.02271).
|
| 10 |
"""
|
|
|
|
| 52 |
"""
|
| 53 |
|
| 54 |
|
| 55 |
+
def write_text(path, text):
|
| 56 |
+
"""Write UTF-8, preserving the file's existing line endings.
|
| 57 |
+
|
| 58 |
+
README.md and CHANGELOG.md are stored CRLF (last regenerated on Windows).
|
| 59 |
+
Writing LF from another OS would rewrite every line, burying the real change
|
| 60 |
+
in a whole-file whitespace diff.
|
| 61 |
+
"""
|
| 62 |
+
text = text.replace("\r\n", "\n")
|
| 63 |
+
if path.exists() and b"\r\n" in path.read_bytes():
|
| 64 |
+
text = text.replace("\n", "\r\n")
|
| 65 |
+
path.write_bytes(text.encode("utf-8"))
|
| 66 |
+
|
| 67 |
+
|
| 68 |
+
def parse_version(text):
|
| 69 |
+
"""'0.2.10' -> (0, 2, 10), so releases order numerically rather than lexically."""
|
| 70 |
+
return tuple(int(part) for part in text.split("."))
|
| 71 |
+
|
| 72 |
+
|
| 73 |
def load_stats(name):
|
| 74 |
f = ROOT / "data" / name / f"{name}.stats.json"
|
| 75 |
return json.loads(f.read_text()) if f.exists() else None
|
| 76 |
|
| 77 |
|
| 78 |
def datasheet(name, cfg, st):
|
| 79 |
+
added = cfg.get("added") or ADDED
|
| 80 |
license_rows = ""
|
| 81 |
licenses = st.get("licenses") or {}
|
| 82 |
if licenses:
|
|
|
|
| 106 |
- **Language:** Polish (pl)
|
| 107 |
- **License:** `{cfg['license']}`
|
| 108 |
- **Created (range):** {cfg['created']}
|
| 109 |
+
- **Added:** {added}
|
| 110 |
|
| 111 |
## Licensing — traceable basis
|
| 112 |
{cfg['traceable']}
|
|
|
|
| 137 |
|
| 138 |
|
| 139 |
def main():
|
| 140 |
+
current = parse_version(VERSION)
|
| 141 |
+
rows, pending, tot_doc, tot_tok, tot_chr = [], [], 0, 0, 0
|
| 142 |
for name, cfg in SOURCES.items():
|
| 143 |
st = load_stats(name)
|
| 144 |
if not st:
|
|
|
|
| 146 |
# Sources with a hand-authored datasheet (rich provenance, per-source add
|
| 147 |
# date) opt out of template regeneration to avoid clobbering it.
|
| 148 |
if not cfg.get("custom_datasheet"):
|
| 149 |
+
write_text(ROOT / "data" / name / f"{name}.md", datasheet(name, cfg, st))
|
| 150 |
+
release = cfg.get("release")
|
| 151 |
+
if release is None or parse_version(release) > current:
|
| 152 |
+
# Built and on main, but not admitted to this release: never counted.
|
| 153 |
+
pending.append((name, cfg, st))
|
| 154 |
+
continue
|
| 155 |
rows.append((name, cfg, st))
|
| 156 |
tot_doc += st["kept"]; tot_tok += st["tokens"]; tot_chr += st["chars"]
|
| 157 |
rows.sort(key=lambda r: -r[2]["tokens"])
|
| 158 |
+
pending.sort(key=lambda r: -r[2]["tokens"])
|
| 159 |
|
| 160 |
tbl = "\n".join(
|
| 161 |
f"| [{n}](data/{n}/{n}.md) | {c['pretty']} | `{c['license']}` | "
|
| 162 |
f"{s['kept']:,} | {s['tokens']/1e6:,.1f}M |"
|
| 163 |
for n, c, s in rows)
|
| 164 |
excl = "\n".join(f"| `{k}` | {v} |" for k, v in EXCLUDED.items())
|
| 165 |
+
unreleased = ""
|
| 166 |
+
if pending:
|
| 167 |
+
unreleased_rows = "\n".join(
|
| 168 |
+
f"| [{n}](data/{n}/{n}.md) | {c['pretty']} | `{c['license']}` | "
|
| 169 |
+
f"{s_['kept']:,} | {s_['tokens']/1e6:,.1f}M |"
|
| 170 |
+
for n, c, s_ in pending)
|
| 171 |
+
unreleased = (
|
| 172 |
+
"\n## Sources on main, not yet released\n"
|
| 173 |
+
f"Built and documented, but **not** part of v{VERSION} and excluded from "
|
| 174 |
+
"every total above. They join a release when a CHANGELOG entry admits "
|
| 175 |
+
"them.\n\n"
|
| 176 |
+
"| source | description | license | documents | tokens |\n"
|
| 177 |
+
"|---|---|---|---:|---:|\n"
|
| 178 |
+
f"{unreleased_rows}\n")
|
| 179 |
phrase_frequency = ""
|
| 180 |
phrase_path = ROOT / "artifacts" / "pattern_frequency_hf_snippet.md"
|
| 181 |
if phrase_path.exists():
|
|
|
|
| 186 |
"\n## Results\n\n"
|
| 187 |
"### Corpus phrase frequency (normalized by tokens)\n\n"
|
| 188 |
"Raw counts and token-normalized shares are regenerated from the "
|
| 189 |
+
"current Parquet files with `src/pattern_frequency_report.py`.\n\n"
|
| 190 |
f"{snippet}"
|
| 191 |
)
|
| 192 |
|
|
|
|
| 249 |
|
| 250 |
#### v0.2.5 — current stable
|
| 251 |
|
| 252 |
+
Added `samorzad_gov_pl`: **72,771 documents / 41,154,506 cl100k-proxy
|
| 253 |
+
tokens** from 246 Polish local public institutions and one central platform
|
| 254 |
+
tenant. The contribution preserves the publishing institution as `author`,
|
| 255 |
+
discovered and fetched article URLs in an attribution sidecar and the verified
|
| 256 |
+
platform-wide CC-BY-SA-4.0 basis.
|
| 257 |
|
| 258 |
#### v0.2.4 — previous stable
|
| 259 |
|
|
|
|
| 358 |
|---|---|---|---:|---:|
|
| 359 |
{tbl}
|
| 360 |
| **total** | | | **{tot_doc:,}** | **{tot_tok/1e6:,.1f}M** |
|
| 361 |
+
{unreleased}
|
| 362 |
## Method
|
| 363 |
Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
|
| 364 |
data. Gates are intentionally minimal (drop short docs, non-Polish, exact duplicates,
|
|
|
|
| 432 |
```
|
| 433 |
{phrase_frequency}
|
| 434 |
"""
|
| 435 |
+
write_text(ROOT / "README.md", readme.rstrip("\n") + "\n")
|
| 436 |
|
| 437 |
# CHANGELOG is append-only release history. Documentation regeneration must
|
| 438 |
# never erase earlier releases or hand-reviewed legal/release notes.
|
|
|
|
| 446 |
f"{tot_doc:,} docs, {tot_tok:,} tokens (tiktoken cl100k proxy).\n\n"
|
| 447 |
)
|
| 448 |
changelog = changelog.replace("# Changelog\n", f"# Changelog\n\n{entry}", 1)
|
| 449 |
+
write_text(changelog_path, changelog)
|
| 450 |
|
| 451 |
+
write_text(
|
| 452 |
+
ROOT / "LICENSE",
|
| 453 |
"Polish DynaWord is released under Creative Commons Attribution-ShareAlike 4.0\n"
|
| 454 |
"International (CC-BY-SA-4.0): https://creativecommons.org/licenses/by-sa/4.0/\n\n"
|
| 455 |
"Per-source upstream licenses and attribution are documented in each\n"
|
| 456 |
"data/<source>/<source>.md datasheet.\n")
|
| 457 |
|
| 458 |
+
print(f"docs written: README + CHANGELOG + LICENSE + "
|
| 459 |
+
f"{len(rows) + len(pending)} datasheets")
|
| 460 |
+
print(f"release v{VERSION}: {len(rows)} sources"
|
| 461 |
+
+ (f" | on main, unreleased: {len(pending)}" if pending else ""))
|
| 462 |
print(f"TOTAL {tot_doc:,} docs | {tot_tok/1e9:.2f}B tok | {tot_chr/1e9:.1f}B chars")
|
| 463 |
|
| 464 |
|
src/sources.py
CHANGED
|
@@ -16,6 +16,8 @@ intermediate aggregator, upstream license/attribution is preserved per source.
|
|
| 16 |
# be rebuilt from direct upstream/export scripts rather than blind SpeakLeash pulls.
|
| 17 |
SOURCES = {
|
| 18 |
"wikipedia": {
|
|
|
|
|
|
|
| 19 |
"speakleash_key": "plwiki",
|
| 20 |
"pretty": "Polish Wikipedia",
|
| 21 |
"license": "CC-BY-SA-3.0",
|
|
@@ -28,6 +30,8 @@ SOURCES = {
|
|
| 28 |
"is_ocr": False,
|
| 29 |
},
|
| 30 |
"wikisource": {
|
|
|
|
|
|
|
| 31 |
"speakleash_key": "plwikisource",
|
| 32 |
"pretty": "Polish Wikisource",
|
| 33 |
"license": "CC-BY-SA-3.0",
|
|
@@ -40,6 +44,8 @@ SOURCES = {
|
|
| 40 |
"is_ocr": False,
|
| 41 |
},
|
| 42 |
"wiktionary_examples": {
|
|
|
|
|
|
|
| 43 |
"file_key": "wiktionary_examples",
|
| 44 |
"pretty": "Polish Wiktionary usage examples",
|
| 45 |
"provenance": "Parsed directly from the official Wikimedia dump "
|
|
@@ -58,6 +64,8 @@ SOURCES = {
|
|
| 58 |
"is_ocr": False,
|
| 59 |
},
|
| 60 |
"eurlex": {
|
|
|
|
|
|
|
| 61 |
"speakleash_key": "eurlex_corpus",
|
| 62 |
"pretty": "EUR-Lex (EU legal acts, Polish)",
|
| 63 |
"license": "CC-BY-4.0",
|
|
@@ -71,6 +79,8 @@ SOURCES = {
|
|
| 71 |
"is_ocr": False,
|
| 72 |
},
|
| 73 |
"parliamentary": {
|
|
|
|
|
|
|
| 74 |
"speakleash_key": "PPC_corpus",
|
| 75 |
"pretty": "Polish Parliamentary Corpus (Sejm/Senat)",
|
| 76 |
"license": "public-domain (official documents)",
|
|
@@ -84,6 +94,9 @@ SOURCES = {
|
|
| 84 |
"is_ocr": False,
|
| 85 |
},
|
| 86 |
"sejm_api": {
|
|
|
|
|
|
|
|
|
|
| 87 |
"file_key": "sejm_api",
|
| 88 |
"pretty": "Sejm API parliamentary speeches (2023 onward)",
|
| 89 |
"license": "public-domain (official documents)",
|
|
@@ -100,6 +113,7 @@ SOURCES = {
|
|
| 100 |
"is_ocr": False,
|
| 101 |
},
|
| 102 |
"sejm_interpellations": {
|
|
|
|
| 103 |
"file_key": "sejm_interpellations",
|
| 104 |
"pretty": "Sejm interpellations and written questions (terms 7–10)",
|
| 105 |
"license": "public-domain (official documents)",
|
|
@@ -119,6 +133,8 @@ SOURCES = {
|
|
| 119 |
"custom_datasheet": True,
|
| 120 |
},
|
| 121 |
"parlamint_pl": {
|
|
|
|
|
|
|
| 122 |
"file_key": "parlamint_pl",
|
| 123 |
"pretty": "ParlaMint-PL (parliamentary debates, 2015-2022)",
|
| 124 |
"license": "CC-BY-4.0",
|
|
@@ -136,6 +152,8 @@ SOURCES = {
|
|
| 136 |
"is_ocr": False,
|
| 137 |
},
|
| 138 |
"wolne_lektury": {
|
|
|
|
|
|
|
| 139 |
"speakleash_key": "wolne_lektury_corpus",
|
| 140 |
"pretty": "Wolne Lektury (school readings)",
|
| 141 |
"license": "CC-BY-SA-4.0 / Wolna Sztuka 1.3",
|
|
@@ -148,6 +166,8 @@ SOURCES = {
|
|
| 148 |
"is_ocr": False,
|
| 149 |
},
|
| 150 |
"wikinews": {
|
|
|
|
|
|
|
| 151 |
"file_key": "plwikinews", # fetched from Wikimedia dumps, NOT via SpeakLeash
|
| 152 |
"pretty": "Polish Wikinews",
|
| 153 |
"license": "CC-BY-2.5",
|
|
@@ -163,6 +183,8 @@ SOURCES = {
|
|
| 163 |
"is_ocr": False,
|
| 164 |
},
|
| 165 |
"wikivoyage": {
|
|
|
|
|
|
|
| 166 |
"file_key": "plwikivoyage",
|
| 167 |
"pretty": "Polish Wikivoyage (travel guides)",
|
| 168 |
"license": "CC-BY-SA-3.0",
|
|
@@ -178,6 +200,8 @@ SOURCES = {
|
|
| 178 |
"is_ocr": False,
|
| 179 |
},
|
| 180 |
"wikibooks": {
|
|
|
|
|
|
|
| 181 |
"file_key": "plwikibooks",
|
| 182 |
"pretty": "Polish Wikibooks (open textbooks)",
|
| 183 |
"license": "CC-BY-SA-3.0",
|
|
@@ -193,6 +217,8 @@ SOURCES = {
|
|
| 193 |
"is_ocr": False,
|
| 194 |
},
|
| 195 |
"wikiquote": {
|
|
|
|
|
|
|
| 196 |
"file_key": "plwikiquote",
|
| 197 |
"pretty": "Polish Wikiquote (quotations)",
|
| 198 |
"license": "CC-BY-SA-3.0",
|
|
@@ -208,6 +234,8 @@ SOURCES = {
|
|
| 208 |
"is_ocr": False,
|
| 209 |
},
|
| 210 |
"eltec_pol": {
|
|
|
|
|
|
|
| 211 |
"file_key": "eltec_pol",
|
| 212 |
"pretty": "ELTeC-pol (European Literary Text Collection, Polish)",
|
| 213 |
"license": "CC-BY-4.0",
|
|
@@ -223,6 +251,8 @@ SOURCES = {
|
|
| 223 |
"is_ocr": False,
|
| 224 |
},
|
| 225 |
"dziennik_ustaw": {
|
|
|
|
|
|
|
| 226 |
"file_key": "dziennik_ustaw", # fetched directly, NOT via SpeakLeash
|
| 227 |
"pretty": "Dziennik Ustaw + Monitor Polski (Polish primary legislation)",
|
| 228 |
"license": "public-domain (official documents)",
|
|
@@ -240,6 +270,8 @@ SOURCES = {
|
|
| 240 |
"is_ocr": False,
|
| 241 |
},
|
| 242 |
"govpl": {
|
|
|
|
|
|
|
| 243 |
"file_key": "govpl", # fetched directly from gov.pl, NOT via SpeakLeash
|
| 244 |
"pretty": "gov.pl — Polish government press releases",
|
| 245 |
"license": "CC-BY-SA-4.0",
|
|
@@ -262,6 +294,8 @@ SOURCES = {
|
|
| 262 |
"is_ocr": False,
|
| 263 |
},
|
| 264 |
"samorzad_gov_pl": {
|
|
|
|
|
|
|
| 265 |
"file_key": "samorzad_gov_pl",
|
| 266 |
"pretty": "samorzad.gov.pl — Polish public-sector institutions",
|
| 267 |
"license": "CC-BY-SA-4.0",
|
|
@@ -285,6 +319,9 @@ SOURCES = {
|
|
| 285 |
"custom_datasheet": True,
|
| 286 |
},
|
| 287 |
"biblioteka_nauki": {
|
|
|
|
|
|
|
|
|
|
| 288 |
"file_key": "biblioteka_nauki_pl_corpus",
|
| 289 |
"pretty": "Biblioteka Nauki",
|
| 290 |
"license": "per-record upstream license",
|
|
@@ -303,6 +340,7 @@ SOURCES = {
|
|
| 303 |
"is_ocr": False,
|
| 304 |
},
|
| 305 |
"europeana": {
|
|
|
|
| 306 |
# Direct rebuild output from fetch_europeana.py, never the raw
|
| 307 |
# SpeakLeash Europeana aggregate.
|
| 308 |
"file_key": "europeana",
|
|
@@ -328,6 +366,9 @@ SOURCES = {
|
|
| 328 |
"legal review confirms their US status.",
|
| 329 |
},
|
| 330 |
"nkjp1m": {
|
|
|
|
|
|
|
|
|
|
| 331 |
"file_key": "nkjp1m",
|
| 332 |
"pretty": "The manually annotated 1-million word subcorpus of the National Corpus of Polish",
|
| 333 |
"license": "CC-BY",
|
|
@@ -340,6 +381,8 @@ SOURCES = {
|
|
| 340 |
"is_ocr": False,
|
| 341 |
},
|
| 342 |
"european_hplt_v3_pl": {
|
|
|
|
|
|
|
| 343 |
"file_key": "european_hplt_v3_pl",
|
| 344 |
"pretty": "HPLT v3.0 Polish (web, top WDS bins 10-5)",
|
| 345 |
"license": "CC0-1.0",
|
|
@@ -371,6 +414,8 @@ SOURCES = {
|
|
| 371 |
"custom_datasheet": True,
|
| 372 |
},
|
| 373 |
"1000_novels": {
|
|
|
|
|
|
|
| 374 |
"speakleash_key": "1000_novels_corpus_CLARIN-PL",
|
| 375 |
"pretty": "1000 Novels Corpus (CLARIN-PL)",
|
| 376 |
"license": "CC-BY-4.0",
|
|
@@ -388,6 +433,8 @@ SOURCES = {
|
|
| 388 |
"custom_datasheet": True,
|
| 389 |
},
|
| 390 |
"global_voices": {
|
|
|
|
|
|
|
| 391 |
# Data contributed directly (PR #7): no in-repo fetch/build script and a
|
| 392 |
# hand-authored datasheet. Registered for the source table + release totals;
|
| 393 |
# the datasheet stays contributor-owned via custom_datasheet.
|
|
@@ -407,6 +454,7 @@ SOURCES = {
|
|
| 407 |
},
|
| 408 |
}
|
| 409 |
|
|
|
|
| 410 |
ADDED = "2026-07-19"
|
| 411 |
EXCLUDED = {
|
| 412 |
"open_subtitles_corpus": "Derivative of copyrighted film/TV dialogue; "
|
|
|
|
| 16 |
# be rebuilt from direct upstream/export scripts rather than blind SpeakLeash pulls.
|
| 17 |
SOURCES = {
|
| 18 |
"wikipedia": {
|
| 19 |
+
"added": "2026-07-02",
|
| 20 |
+
"release": "0.2.0",
|
| 21 |
"speakleash_key": "plwiki",
|
| 22 |
"pretty": "Polish Wikipedia",
|
| 23 |
"license": "CC-BY-SA-3.0",
|
|
|
|
| 30 |
"is_ocr": False,
|
| 31 |
},
|
| 32 |
"wikisource": {
|
| 33 |
+
"added": "2026-07-02",
|
| 34 |
+
"release": "0.2.0",
|
| 35 |
"speakleash_key": "plwikisource",
|
| 36 |
"pretty": "Polish Wikisource",
|
| 37 |
"license": "CC-BY-SA-3.0",
|
|
|
|
| 44 |
"is_ocr": False,
|
| 45 |
},
|
| 46 |
"wiktionary_examples": {
|
| 47 |
+
"added": "2026-07-19",
|
| 48 |
+
"release": None,
|
| 49 |
"file_key": "wiktionary_examples",
|
| 50 |
"pretty": "Polish Wiktionary usage examples",
|
| 51 |
"provenance": "Parsed directly from the official Wikimedia dump "
|
|
|
|
| 64 |
"is_ocr": False,
|
| 65 |
},
|
| 66 |
"eurlex": {
|
| 67 |
+
"added": "2026-07-02",
|
| 68 |
+
"release": "0.2.0",
|
| 69 |
"speakleash_key": "eurlex_corpus",
|
| 70 |
"pretty": "EUR-Lex (EU legal acts, Polish)",
|
| 71 |
"license": "CC-BY-4.0",
|
|
|
|
| 79 |
"is_ocr": False,
|
| 80 |
},
|
| 81 |
"parliamentary": {
|
| 82 |
+
"added": "2026-07-02",
|
| 83 |
+
"release": "0.2.0",
|
| 84 |
"speakleash_key": "PPC_corpus",
|
| 85 |
"pretty": "Polish Parliamentary Corpus (Sejm/Senat)",
|
| 86 |
"license": "public-domain (official documents)",
|
|
|
|
| 94 |
"is_ocr": False,
|
| 95 |
},
|
| 96 |
"sejm_api": {
|
| 97 |
+
"added": "2026-08-22",
|
| 98 |
+
"custom_datasheet": True,
|
| 99 |
+
"release": None,
|
| 100 |
"file_key": "sejm_api",
|
| 101 |
"pretty": "Sejm API parliamentary speeches (2023 onward)",
|
| 102 |
"license": "public-domain (official documents)",
|
|
|
|
| 113 |
"is_ocr": False,
|
| 114 |
},
|
| 115 |
"sejm_interpellations": {
|
| 116 |
+
"release": None,
|
| 117 |
"file_key": "sejm_interpellations",
|
| 118 |
"pretty": "Sejm interpellations and written questions (terms 7–10)",
|
| 119 |
"license": "public-domain (official documents)",
|
|
|
|
| 133 |
"custom_datasheet": True,
|
| 134 |
},
|
| 135 |
"parlamint_pl": {
|
| 136 |
+
"added": "2026-07-19",
|
| 137 |
+
"release": None,
|
| 138 |
"file_key": "parlamint_pl",
|
| 139 |
"pretty": "ParlaMint-PL (parliamentary debates, 2015-2022)",
|
| 140 |
"license": "CC-BY-4.0",
|
|
|
|
| 152 |
"is_ocr": False,
|
| 153 |
},
|
| 154 |
"wolne_lektury": {
|
| 155 |
+
"added": "2026-07-02",
|
| 156 |
+
"release": "0.2.0",
|
| 157 |
"speakleash_key": "wolne_lektury_corpus",
|
| 158 |
"pretty": "Wolne Lektury (school readings)",
|
| 159 |
"license": "CC-BY-SA-4.0 / Wolna Sztuka 1.3",
|
|
|
|
| 166 |
"is_ocr": False,
|
| 167 |
},
|
| 168 |
"wikinews": {
|
| 169 |
+
"added": "2026-07-02",
|
| 170 |
+
"release": "0.2.0",
|
| 171 |
"file_key": "plwikinews", # fetched from Wikimedia dumps, NOT via SpeakLeash
|
| 172 |
"pretty": "Polish Wikinews",
|
| 173 |
"license": "CC-BY-2.5",
|
|
|
|
| 183 |
"is_ocr": False,
|
| 184 |
},
|
| 185 |
"wikivoyage": {
|
| 186 |
+
"added": "2026-07-02",
|
| 187 |
+
"release": "0.2.0",
|
| 188 |
"file_key": "plwikivoyage",
|
| 189 |
"pretty": "Polish Wikivoyage (travel guides)",
|
| 190 |
"license": "CC-BY-SA-3.0",
|
|
|
|
| 200 |
"is_ocr": False,
|
| 201 |
},
|
| 202 |
"wikibooks": {
|
| 203 |
+
"added": "2026-07-02",
|
| 204 |
+
"release": "0.2.0",
|
| 205 |
"file_key": "plwikibooks",
|
| 206 |
"pretty": "Polish Wikibooks (open textbooks)",
|
| 207 |
"license": "CC-BY-SA-3.0",
|
|
|
|
| 217 |
"is_ocr": False,
|
| 218 |
},
|
| 219 |
"wikiquote": {
|
| 220 |
+
"added": "2026-07-02",
|
| 221 |
+
"release": "0.2.0",
|
| 222 |
"file_key": "plwikiquote",
|
| 223 |
"pretty": "Polish Wikiquote (quotations)",
|
| 224 |
"license": "CC-BY-SA-3.0",
|
|
|
|
| 234 |
"is_ocr": False,
|
| 235 |
},
|
| 236 |
"eltec_pol": {
|
| 237 |
+
"added": "2026-07-02",
|
| 238 |
+
"release": "0.2.0",
|
| 239 |
"file_key": "eltec_pol",
|
| 240 |
"pretty": "ELTeC-pol (European Literary Text Collection, Polish)",
|
| 241 |
"license": "CC-BY-4.0",
|
|
|
|
| 251 |
"is_ocr": False,
|
| 252 |
},
|
| 253 |
"dziennik_ustaw": {
|
| 254 |
+
"added": "2026-07-02",
|
| 255 |
+
"release": "0.2.0",
|
| 256 |
"file_key": "dziennik_ustaw", # fetched directly, NOT via SpeakLeash
|
| 257 |
"pretty": "Dziennik Ustaw + Monitor Polski (Polish primary legislation)",
|
| 258 |
"license": "public-domain (official documents)",
|
|
|
|
| 270 |
"is_ocr": False,
|
| 271 |
},
|
| 272 |
"govpl": {
|
| 273 |
+
"added": "2026-07-19",
|
| 274 |
+
"release": "0.2.2",
|
| 275 |
"file_key": "govpl", # fetched directly from gov.pl, NOT via SpeakLeash
|
| 276 |
"pretty": "gov.pl — Polish government press releases",
|
| 277 |
"license": "CC-BY-SA-4.0",
|
|
|
|
| 294 |
"is_ocr": False,
|
| 295 |
},
|
| 296 |
"samorzad_gov_pl": {
|
| 297 |
+
"added": "2026-08-14",
|
| 298 |
+
"release": "0.2.5",
|
| 299 |
"file_key": "samorzad_gov_pl",
|
| 300 |
"pretty": "samorzad.gov.pl — Polish public-sector institutions",
|
| 301 |
"license": "CC-BY-SA-4.0",
|
|
|
|
| 319 |
"custom_datasheet": True,
|
| 320 |
},
|
| 321 |
"biblioteka_nauki": {
|
| 322 |
+
"added": "2026-06-27",
|
| 323 |
+
"custom_datasheet": True,
|
| 324 |
+
"release": "0.2.4",
|
| 325 |
"file_key": "biblioteka_nauki_pl_corpus",
|
| 326 |
"pretty": "Biblioteka Nauki",
|
| 327 |
"license": "per-record upstream license",
|
|
|
|
| 340 |
"is_ocr": False,
|
| 341 |
},
|
| 342 |
"europeana": {
|
| 343 |
+
"release": None,
|
| 344 |
# Direct rebuild output from fetch_europeana.py, never the raw
|
| 345 |
# SpeakLeash Europeana aggregate.
|
| 346 |
"file_key": "europeana",
|
|
|
|
| 366 |
"legal review confirms their US status.",
|
| 367 |
},
|
| 368 |
"nkjp1m": {
|
| 369 |
+
"added": "2026-06-15",
|
| 370 |
+
"custom_datasheet": True,
|
| 371 |
+
"release": "0.2.3",
|
| 372 |
"file_key": "nkjp1m",
|
| 373 |
"pretty": "The manually annotated 1-million word subcorpus of the National Corpus of Polish",
|
| 374 |
"license": "CC-BY",
|
|
|
|
| 381 |
"is_ocr": False,
|
| 382 |
},
|
| 383 |
"european_hplt_v3_pl": {
|
| 384 |
+
"added": "2026-07-14",
|
| 385 |
+
"release": "0.2.3",
|
| 386 |
"file_key": "european_hplt_v3_pl",
|
| 387 |
"pretty": "HPLT v3.0 Polish (web, top WDS bins 10-5)",
|
| 388 |
"license": "CC0-1.0",
|
|
|
|
| 414 |
"custom_datasheet": True,
|
| 415 |
},
|
| 416 |
"1000_novels": {
|
| 417 |
+
"added": "2026-07-02",
|
| 418 |
+
"release": "0.2.1",
|
| 419 |
"speakleash_key": "1000_novels_corpus_CLARIN-PL",
|
| 420 |
"pretty": "1000 Novels Corpus (CLARIN-PL)",
|
| 421 |
"license": "CC-BY-4.0",
|
|
|
|
| 433 |
"custom_datasheet": True,
|
| 434 |
},
|
| 435 |
"global_voices": {
|
| 436 |
+
"added": "2026-07-15",
|
| 437 |
+
"release": "0.2.3",
|
| 438 |
# Data contributed directly (PR #7): no in-repo fetch/build script and a
|
| 439 |
# hand-authored datasheet. Registered for the source table + release totals;
|
| 440 |
# the datasheet stays contributor-owned via custom_datasheet.
|
|
|
|
| 454 |
},
|
| 455 |
}
|
| 456 |
|
| 457 |
+
# Fallback stamp for a source with no recorded "added" date of its own.
|
| 458 |
ADDED = "2026-07-19"
|
| 459 |
EXCLUDED = {
|
| 460 |
"open_subtitles_corpus": "Derivative of copyrighted film/TV dialogue; "
|
src/test_verify_paper_claims.py
ADDED
|
@@ -0,0 +1,190 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Self-check for src/verify_paper_claims.py against a synthetic mini-repo."""
|
| 3 |
+
|
| 4 |
+
import json
|
| 5 |
+
import shutil
|
| 6 |
+
import sys
|
| 7 |
+
import tempfile
|
| 8 |
+
from pathlib import Path
|
| 9 |
+
|
| 10 |
+
sys.path.insert(0, str(Path(__file__).resolve().parent))
|
| 11 |
+
import verify_paper_claims as vpc # noqa: E402
|
| 12 |
+
|
| 13 |
+
TEX = r"""
|
| 14 |
+
\begin{figure}
|
| 15 |
+
\addplot[thick] coordinates {(v0.1.0,1.00)(v0.1.1,3.00)};
|
| 16 |
+
\caption{grew from 1.00 billion tokens and 1 sources (v0.1.0).}
|
| 17 |
+
\end{figure}
|
| 18 |
+
\begin{abstract}
|
| 19 |
+
Release v0.1.1 comprises \num{300} documents and \num{3.00}~billion tokens
|
| 20 |
+
from 2 sources.
|
| 21 |
+
\end{abstract}
|
| 22 |
+
\begin{table}
|
| 23 |
+
\begin{tabular}{lrr}
|
| 24 |
+
Version & Tokens (billions) & Sources \\
|
| 25 |
+
v0.1.0 & 1.00 & 1 \\
|
| 26 |
+
v0.1.1 & 3.00 & 2 \\
|
| 27 |
+
\end{tabular}
|
| 28 |
+
\label{tab:s04-wersje}
|
| 29 |
+
\end{table}
|
| 30 |
+
\begin{table}
|
| 31 |
+
\label{tab:s05-zrodla}
|
| 32 |
+
\begin{tabular}{llr}
|
| 33 |
+
Source & License & Tokens [M] \\
|
| 34 |
+
\texttt{alpha} & CC0 & 2000.0 \\
|
| 35 |
+
\texttt{beta\_two} & CC-BY-SA-4.0 & 1000.0 \\
|
| 36 |
+
\textbf{Total} & CC-BY-SA-4.0 & \textbf{$\approx$3000} \\
|
| 37 |
+
\end{tabular}
|
| 38 |
+
\end{table}
|
| 39 |
+
\begin{table}
|
| 40 |
+
\label{tab:s10-licencje}
|
| 41 |
+
\begin{tabular}{llr}
|
| 42 |
+
License family & Representative sources & Tokens [billions] \\
|
| 43 |
+
CC0 & alpha & $2.00$ \\
|
| 44 |
+
CC-BY-SA & beta two & $1.00$ \\
|
| 45 |
+
Total & & $3.00$ \\
|
| 46 |
+
\end{tabular}
|
| 47 |
+
\end{table}
|
| 48 |
+
This work is pinned to commit \texttt{deadbeef}.
|
| 49 |
+
"""
|
| 50 |
+
|
| 51 |
+
CHANGELOG = """# Changelog
|
| 52 |
+
|
| 53 |
+
## v0.1.1 (2026-01-02)
|
| 54 |
+
|
| 55 |
+
- Added `beta_two`.
|
| 56 |
+
- Totals: 2 sources, 300 docs, 3.00B tokens.
|
| 57 |
+
"""
|
| 58 |
+
|
| 59 |
+
README = "| `v0.1.0` | previous stable | 100 | 1.00B | Baseline from 1 open/official sources. |\n"
|
| 60 |
+
|
| 61 |
+
GROWTH = """VERSIONS = ["v0.1.0", "v0.1.1"]
|
| 62 |
+
DOCUMENTS = [100, 300]
|
| 63 |
+
TOKENS = [1_000_000_000, 3_000_000_000]
|
| 64 |
+
"""
|
| 65 |
+
|
| 66 |
+
STATS = {"alpha": (200, 2_000_000_000), "beta_two": (100, 1_000_000_000)}
|
| 67 |
+
|
| 68 |
+
SOURCES = """SOURCES = {
|
| 69 |
+
"alpha": {"license": "CC0-1.0", "license_spdx": "CC0-1.0", "release": "0.1.0"},
|
| 70 |
+
"beta_two": {"license": "CC-BY-SA-4.0", "license_spdx": "CC-BY-SA-4.0", "release": "0.1.1"},
|
| 71 |
+
}
|
| 72 |
+
"""
|
| 73 |
+
|
| 74 |
+
|
| 75 |
+
def registry(root: Path) -> dict:
|
| 76 |
+
return vpc.load_registry(root)
|
| 77 |
+
|
| 78 |
+
|
| 79 |
+
def edit(path: Path, old: str, new: str) -> None:
|
| 80 |
+
text = path.read_text()
|
| 81 |
+
assert old in text, f"fixture no longer contains {old!r}"
|
| 82 |
+
path.write_text(text.replace(old, new))
|
| 83 |
+
|
| 84 |
+
|
| 85 |
+
def build(root: Path) -> None:
|
| 86 |
+
(root / "src").mkdir(parents=True)
|
| 87 |
+
(root / "CHANGELOG.md").write_text(CHANGELOG)
|
| 88 |
+
(root / "README.md").write_text(README)
|
| 89 |
+
(root / "src" / "plot_version_growth.py").write_text(GROWTH)
|
| 90 |
+
(root / "src" / "sources.py").write_text(SOURCES)
|
| 91 |
+
for name, (docs, tokens) in STATS.items():
|
| 92 |
+
directory = root / "data" / name
|
| 93 |
+
directory.mkdir(parents=True)
|
| 94 |
+
(directory / f"{name}.stats.json").write_text(
|
| 95 |
+
json.dumps({"kept": docs, "tokens": tokens, "license": "CC0"})
|
| 96 |
+
)
|
| 97 |
+
(directory / f"{name}.md").write_text(
|
| 98 |
+
f"# {name}\n- **License:** `{registry(root)[name]['license']}`\n"
|
| 99 |
+
)
|
| 100 |
+
(root / "paper.tex").write_text(TEX)
|
| 101 |
+
|
| 102 |
+
|
| 103 |
+
def run(root: Path) -> int:
|
| 104 |
+
return vpc.main(["--tex", str(root / "paper.tex"), "--root", str(root), "--ref", "worktree"])
|
| 105 |
+
|
| 106 |
+
|
| 107 |
+
def scenario(mutate=None) -> int:
|
| 108 |
+
root = Path(tempfile.mkdtemp())
|
| 109 |
+
try:
|
| 110 |
+
build(root)
|
| 111 |
+
if mutate:
|
| 112 |
+
mutate(root)
|
| 113 |
+
return run(root)
|
| 114 |
+
finally:
|
| 115 |
+
shutil.rmtree(root)
|
| 116 |
+
|
| 117 |
+
|
| 118 |
+
def drop_token_digit(root: Path) -> None:
|
| 119 |
+
path = root / "paper.tex"
|
| 120 |
+
path.write_text(path.read_text().replace("2000.0", "2500.0"))
|
| 121 |
+
|
| 122 |
+
|
| 123 |
+
def wrong_source_count(root: Path) -> None:
|
| 124 |
+
path = root / "paper.tex"
|
| 125 |
+
path.write_text(path.read_text().replace("v0.1.1 & 3.00 & 2", "v0.1.1 & 3.00 & ---"))
|
| 126 |
+
|
| 127 |
+
|
| 128 |
+
def add_source(root: Path, release: str | None) -> None:
|
| 129 |
+
directory = root / "data" / "gamma"
|
| 130 |
+
directory.mkdir(parents=True)
|
| 131 |
+
(directory / "gamma.stats.json").write_text(json.dumps({"kept": 1, "tokens": 1, "license": "CC0"}))
|
| 132 |
+
(directory / "gamma.md").write_text("# gamma\n- **License:** `CC0-1.0`\n")
|
| 133 |
+
edit(root / "src" / "sources.py", "}\n",
|
| 134 |
+
f' "gamma": {{"license": "CC0-1.0", "release": {release!r}}},\n}}\n')
|
| 135 |
+
|
| 136 |
+
|
| 137 |
+
def unlisted_source(root: Path) -> None:
|
| 138 |
+
add_source(root, "0.1.1")
|
| 139 |
+
|
| 140 |
+
|
| 141 |
+
def unreleased_source(root: Path) -> None:
|
| 142 |
+
add_source(root, None)
|
| 143 |
+
|
| 144 |
+
|
| 145 |
+
def wrong_license(root: Path) -> None:
|
| 146 |
+
edit(root / "paper.tex", "CC-BY-SA-4.0 & 1000.0", "CC-BY-4.0 & 1000.0")
|
| 147 |
+
|
| 148 |
+
|
| 149 |
+
def datasheet_license_drift(root: Path) -> None:
|
| 150 |
+
edit(root / "data" / "alpha" / "alpha.md", "CC0-1.0", "CC-BY-4.0")
|
| 151 |
+
|
| 152 |
+
|
| 153 |
+
def undefined_footnote(root: Path) -> None:
|
| 154 |
+
edit(root / "paper.tex", "& CC0 &", r"& CC0$^{\dagger}$ &")
|
| 155 |
+
|
| 156 |
+
|
| 157 |
+
def family_value_drift(root: Path) -> None:
|
| 158 |
+
edit(root / "paper.tex", "& beta two & $1.00$", "& beta two & $1.50$")
|
| 159 |
+
|
| 160 |
+
|
| 161 |
+
def incomplete_partition(root: Path) -> None:
|
| 162 |
+
edit(root / "paper.tex", "CC-BY-SA & beta two & $1.00$ \\\\\n", "")
|
| 163 |
+
|
| 164 |
+
|
| 165 |
+
def unfootnoted_dual_license(root: Path) -> None:
|
| 166 |
+
"""alpha's content is CC0 but its platform redistributes under CC-BY-4.0."""
|
| 167 |
+
edit(root / "src" / "sources.py",
|
| 168 |
+
'"alpha": {"license": "CC0-1.0", "license_spdx": "CC0-1.0"',
|
| 169 |
+
'"alpha": {"license": "CC0-1.0", "license_spdx": "CC-BY-4.0"')
|
| 170 |
+
|
| 171 |
+
|
| 172 |
+
def stale_growth_series(root: Path) -> None:
|
| 173 |
+
path = root / "src" / "plot_version_growth.py"
|
| 174 |
+
path.write_text(path.read_text().replace("3_000_000_000", "3_100_000_000"))
|
| 175 |
+
|
| 176 |
+
|
| 177 |
+
if __name__ == "__main__":
|
| 178 |
+
assert scenario() == 0, "consistent fixture must pass"
|
| 179 |
+
assert scenario(drop_token_digit) == 1, "per-source token drift must fail"
|
| 180 |
+
assert scenario(wrong_source_count) == 1, "carried-over '---' source count must fail"
|
| 181 |
+
assert scenario(unlisted_source) == 1, "source present in repo but not in Table 3 must fail"
|
| 182 |
+
assert scenario(stale_growth_series) == 1, "growth series vs composition drift must fail"
|
| 183 |
+
assert scenario(unreleased_source) == 0, "a source with release=None is not part of the release"
|
| 184 |
+
assert scenario(wrong_license) == 1, "license in Table 3 vs registry must fail"
|
| 185 |
+
assert scenario(datasheet_license_drift) == 1, "datasheet vs registry license drift must fail"
|
| 186 |
+
assert scenario(undefined_footnote) == 1, "a footnote marker with no definition must fail"
|
| 187 |
+
assert scenario(unfootnoted_dual_license) == 1, "an unfootnoted dual license must fail"
|
| 188 |
+
assert scenario(family_value_drift) == 1, "a wrong license-family subtotal must fail"
|
| 189 |
+
assert scenario(incomplete_partition) == 1, "a license family with no row must fail"
|
| 190 |
+
print("ok")
|
src/verify_paper_claims.py
ADDED
|
@@ -0,0 +1,637 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
#!/usr/bin/env python3
|
| 2 |
+
"""Verify the paper's numeric claims against repository ground truth.
|
| 3 |
+
|
| 4 |
+
Two checks, matching the two things the paper asserts numerically:
|
| 5 |
+
|
| 6 |
+
1. Release trajectory (Table `tab:s04-wersje` + the hero figure):
|
| 7 |
+
tokens and source counts for v0.2.0 -> v0.2.5.
|
| 8 |
+
Ground truth: `src/plot_version_growth.py` literals (exact docs/tokens)
|
| 9 |
+
and CHANGELOG.md / README.md (source counts per release).
|
| 10 |
+
|
| 11 |
+
2. Corpus composition (Table `tab:s05-zrodla` + the inline totals):
|
| 12 |
+
per-source token counts, the source set, and the corpus totals.
|
| 13 |
+
Ground truth: `data/*/*.stats.json` at a git ref.
|
| 14 |
+
|
| 15 |
+
Usage:
|
| 16 |
+
python src/verify_paper_claims.py # ref = pin found in the .tex
|
| 17 |
+
python src/verify_paper_claims.py --ref worktree # current working tree
|
| 18 |
+
python src/verify_paper_claims.py --tex paper.tex --ref HEAD
|
| 19 |
+
"""
|
| 20 |
+
|
| 21 |
+
from __future__ import annotations
|
| 22 |
+
|
| 23 |
+
import argparse
|
| 24 |
+
import ast
|
| 25 |
+
import json
|
| 26 |
+
import re
|
| 27 |
+
import subprocess
|
| 28 |
+
import sys
|
| 29 |
+
from pathlib import Path
|
| 30 |
+
|
| 31 |
+
ROOT = Path(__file__).resolve().parent.parent
|
| 32 |
+
|
| 33 |
+
# Printed precision in the paper sets the comparison tolerance.
|
| 34 |
+
TOL_BILLIONS = 0.005 # Table 2 / hero figure: 2 decimal places
|
| 35 |
+
TOL_MILLIONS = 0.05 # Table 3: 1 decimal place
|
| 36 |
+
|
| 37 |
+
OK, FAIL, WARN = "OK", "FAIL", "WARN"
|
| 38 |
+
|
| 39 |
+
MARKER = re.compile(r"\$\^\{([^}]*)\}\$")
|
| 40 |
+
|
| 41 |
+
|
| 42 |
+
class Report:
|
| 43 |
+
def __init__(self) -> None:
|
| 44 |
+
self.rows: list[tuple[str, str, str]] = []
|
| 45 |
+
|
| 46 |
+
def add(self, status: str, label: str, detail: str = "") -> None:
|
| 47 |
+
self.rows.append((status, label, detail))
|
| 48 |
+
|
| 49 |
+
def num(self, label: str, claimed, actual, tol: float = 0.0) -> None:
|
| 50 |
+
if claimed is None:
|
| 51 |
+
self.add(WARN, label, "not found in .tex")
|
| 52 |
+
elif abs(claimed - actual) <= tol:
|
| 53 |
+
self.add(OK, label, f"{claimed} == {actual}")
|
| 54 |
+
else:
|
| 55 |
+
self.add(FAIL, label, f"paper {claimed} != repo {actual}")
|
| 56 |
+
|
| 57 |
+
def section(self, title: str) -> None:
|
| 58 |
+
self.rows.append(("", title, ""))
|
| 59 |
+
|
| 60 |
+
def render(self) -> int:
|
| 61 |
+
width = max(len(label) for _, label, _ in self.rows) + 2
|
| 62 |
+
for status, label, detail in self.rows:
|
| 63 |
+
if not status:
|
| 64 |
+
print(f"\n=== {label} ===")
|
| 65 |
+
else:
|
| 66 |
+
print(f" [{status:<4}] {label:<{width}} {detail}")
|
| 67 |
+
failures = sum(1 for s, _, _ in self.rows if s == FAIL)
|
| 68 |
+
warnings = sum(1 for s, _, _ in self.rows if s == WARN)
|
| 69 |
+
print(f"\n{failures} failure(s), {warnings} warning(s)")
|
| 70 |
+
return 1 if failures else 0
|
| 71 |
+
|
| 72 |
+
|
| 73 |
+
# --------------------------------------------------------------------------
|
| 74 |
+
# .tex parsing
|
| 75 |
+
# --------------------------------------------------------------------------
|
| 76 |
+
|
| 77 |
+
def _unescape(name: str) -> str:
|
| 78 |
+
return name.replace("\\_", "_").replace("\\&", "&")
|
| 79 |
+
|
| 80 |
+
|
| 81 |
+
def _table_body(tex: str, label: str) -> str:
|
| 82 |
+
"""Return the body of the table environment carrying \\label{label}."""
|
| 83 |
+
for block in re.findall(r"\\begin\{table\}.*?\\end\{table\}", tex, re.S):
|
| 84 |
+
if f"\\label{{{label}}}" in block:
|
| 85 |
+
return block
|
| 86 |
+
raise LookupError(f"no table with \\label{{{label}}} in the .tex")
|
| 87 |
+
|
| 88 |
+
|
| 89 |
+
def parse_hero(tex: str) -> dict[str, float]:
|
| 90 |
+
"""Version -> tokens (billions) from the pgfplots hero figure coordinates."""
|
| 91 |
+
coords = re.search(r"\\addplot\[[^\]]*\]\s*coordinates\s*\{(.*?)\}", tex, re.S)
|
| 92 |
+
if not coords:
|
| 93 |
+
return {}
|
| 94 |
+
return {v: float(t) for v, t in re.findall(r"\((v[\d.]+),\s*([\d.]+)\)", coords.group(1))}
|
| 95 |
+
|
| 96 |
+
|
| 97 |
+
def parse_trajectory_table(tex: str) -> dict[str, tuple[float, str]]:
|
| 98 |
+
"""Version -> (tokens in billions, source-count cell verbatim)."""
|
| 99 |
+
body = _table_body(tex, "tab:s04-wersje")
|
| 100 |
+
rows = re.findall(r"^\s*(v[\d.]+)\s*&\s*([\d.]+)\s*&\s*(.+?)\s*\\\\", body, re.M)
|
| 101 |
+
return {v: (float(tok), src.strip()) for v, tok, src in rows}
|
| 102 |
+
|
| 103 |
+
|
| 104 |
+
def parse_source_table(tex: str) -> tuple[dict[str, float], float | None]:
|
| 105 |
+
"""(source name -> tokens in millions, declared total in millions)."""
|
| 106 |
+
body = _table_body(tex, "tab:s05-zrodla")
|
| 107 |
+
sources = {
|
| 108 |
+
_unescape(name): float(tok)
|
| 109 |
+
for name, tok in re.findall(
|
| 110 |
+
r"\\texttt\{([^}]+)\}\s*&.*?&\s*([\d.]+)\s*\\\\", body
|
| 111 |
+
)
|
| 112 |
+
}
|
| 113 |
+
total_row = re.search(r"\\textbf\{Total\}.*?&.*?&\s*(.+?)\s*\\\\", body, re.S)
|
| 114 |
+
total = None
|
| 115 |
+
if total_row:
|
| 116 |
+
digits = re.search(r"([\d.]+)", total_row.group(1))
|
| 117 |
+
if digits:
|
| 118 |
+
total = float(digits.group(1))
|
| 119 |
+
return sources, total
|
| 120 |
+
|
| 121 |
+
|
| 122 |
+
def parse_license_rows(tex: str) -> dict[str, tuple[str, set[str]]]:
|
| 123 |
+
"""Table 3 source rows -> (license cell without its markers, marker names)."""
|
| 124 |
+
body = _table_body(tex, "tab:s05-zrodla").split(r"\begin{tabular}")[-1]
|
| 125 |
+
rows = {}
|
| 126 |
+
for name, cell in re.findall(r"\\texttt\{([^}]+)\}\s*&(.*?)&\s*[\d.]+\s*\\\\", body):
|
| 127 |
+
rows[_unescape(name)] = (MARKER.sub("", cell).strip(), _markers(cell))
|
| 128 |
+
return rows
|
| 129 |
+
|
| 130 |
+
|
| 131 |
+
def parse_footnote_defs(tex: str) -> set[str]:
|
| 132 |
+
"""Markers the Table 3 caption actually explains (a definition is followed by ~)."""
|
| 133 |
+
caption = _table_body(tex, "tab:s05-zrodla").split(r"\begin{tabular}")[0]
|
| 134 |
+
return {m for cell in re.findall(r"(\$\^\{[^}]*\}\$)~", caption) for m in _markers(cell)}
|
| 135 |
+
|
| 136 |
+
|
| 137 |
+
def parse_total_license(tex: str) -> str | None:
|
| 138 |
+
body = _table_body(tex, "tab:s05-zrodla")
|
| 139 |
+
row = re.search(r"\\textbf\{Total\}\s*&\s*(.*?)\s*&", body)
|
| 140 |
+
return MARKER.sub("", row.group(1)).strip() if row else None
|
| 141 |
+
|
| 142 |
+
|
| 143 |
+
def _markers(cell: str) -> set[str]:
|
| 144 |
+
return {m for group in MARKER.findall(cell) for m in re.findall(r"\\([a-zA-Z]+)", group)}
|
| 145 |
+
|
| 146 |
+
|
| 147 |
+
def parse_family_table(tex: str) -> tuple[dict[str, float], float | None]:
|
| 148 |
+
"""Table 5 -> (license family -> tokens in billions, declared total)."""
|
| 149 |
+
body = _table_body(tex, "tab:s10-licencje").split(r"\begin{tabular}")[-1]
|
| 150 |
+
families, total = {}, None
|
| 151 |
+
for label, value in re.findall(r"^\s*([^&\\]+?)\s*&[^&]*&\s*\$?([\d.]+)\$?\s*\\\\",
|
| 152 |
+
body, re.M):
|
| 153 |
+
if label.strip().lower() == "total":
|
| 154 |
+
total = float(value)
|
| 155 |
+
elif license_family(label):
|
| 156 |
+
families[license_family(label)] = float(value)
|
| 157 |
+
return families, total
|
| 158 |
+
|
| 159 |
+
|
| 160 |
+
def license_family(text: str) -> str | None:
|
| 161 |
+
"""The Table 5 partition: like norm_license, but CC0 and public domain split.
|
| 162 |
+
|
| 163 |
+
norm_license folds them together because the source table prints either for
|
| 164 |
+
the same source; Table 5 reports them as separate families, so a source's
|
| 165 |
+
own wording decides which side it lands on.
|
| 166 |
+
"""
|
| 167 |
+
t = re.sub(r"[`_]", " ", text).upper()
|
| 168 |
+
if re.search(r"PER-RECORD|PER-DOCUMENT|MIXED", t):
|
| 169 |
+
return "Per-record (mixed)"
|
| 170 |
+
if re.search(r"PUBLIC[- ]DOMAIN", t):
|
| 171 |
+
return "Public domain"
|
| 172 |
+
if "CC0" in t:
|
| 173 |
+
return "CC0"
|
| 174 |
+
if re.search(r"CC[- ]BY[- ]SA", t):
|
| 175 |
+
return "CC-BY-SA"
|
| 176 |
+
if re.search(r"CC[- ]BY", t):
|
| 177 |
+
return "CC-BY"
|
| 178 |
+
return None
|
| 179 |
+
|
| 180 |
+
|
| 181 |
+
def parse_inline_claims(tex: str, current: str | None = None) -> dict:
|
| 182 |
+
"""Claims the prose/captions make about the *current release* totals.
|
| 183 |
+
|
| 184 |
+
A claim is skipped when its surrounding sentence scopes it to a different
|
| 185 |
+
release or to a partition subtotal, so only whole-release assertions are
|
| 186 |
+
compared against the current release's ground truth.
|
| 187 |
+
"""
|
| 188 |
+
CONTEXT = 160
|
| 189 |
+
SUBTOTAL = re.compile(r"share-alike|partitions?\b")
|
| 190 |
+
VERSION = re.compile(r"v\d+\.\d+\.\d+")
|
| 191 |
+
|
| 192 |
+
def sweep(pattern: str, cast):
|
| 193 |
+
out = set()
|
| 194 |
+
for match in re.finditer(pattern, tex):
|
| 195 |
+
window = tex[max(0, match.start() - CONTEXT):match.end() + CONTEXT]
|
| 196 |
+
if SUBTOTAL.search(window):
|
| 197 |
+
continue
|
| 198 |
+
if any(v != current for v in VERSION.findall(window)):
|
| 199 |
+
continue
|
| 200 |
+
raw = next(g for g in match.groups() if g)
|
| 201 |
+
out.add(cast(raw))
|
| 202 |
+
return out
|
| 203 |
+
|
| 204 |
+
docs = sweep(
|
| 205 |
+
r"(?:\\num\{(\d+)\}|(\d[\d\\,]*\d))\s*~?\s*documents",
|
| 206 |
+
lambda raw: int(re.sub(r"[^\d]", "", raw)),
|
| 207 |
+
)
|
| 208 |
+
# ponytail: floor keeps small unrelated "N documents" counts (held-out set
|
| 209 |
+
# sizes) out; scope by section if the paper grows other six-figure claims.
|
| 210 |
+
docs = {value for value in docs if value >= 100_000}
|
| 211 |
+
|
| 212 |
+
tokens = sweep(
|
| 213 |
+
r"(?:\\num\{([\d.]+)\}|\$([\d.]+)\$|([\d.]+))\s*~?\s*billion\s+tokens",
|
| 214 |
+
float,
|
| 215 |
+
)
|
| 216 |
+
sources = sweep(r"\b(\d+)\s+sources\b", int)
|
| 217 |
+
|
| 218 |
+
pin = re.search(r"commit\s+\\texttt\{([0-9a-f]{7,40})\}", tex)
|
| 219 |
+
return {
|
| 220 |
+
"documents": docs,
|
| 221 |
+
"tokens_billions": tokens,
|
| 222 |
+
"sources": sources,
|
| 223 |
+
"pin": pin.group(1) if pin else None,
|
| 224 |
+
}
|
| 225 |
+
|
| 226 |
+
|
| 227 |
+
# --------------------------------------------------------------------------
|
| 228 |
+
# Ground truth
|
| 229 |
+
# --------------------------------------------------------------------------
|
| 230 |
+
|
| 231 |
+
def load_growth_series(root: Path) -> dict[str, tuple[int, int]]:
|
| 232 |
+
"""Version -> (documents, tokens) from src/plot_version_growth.py literals."""
|
| 233 |
+
path = root / "src" / "plot_version_growth.py"
|
| 234 |
+
tree = ast.parse(path.read_text(encoding="utf-8"))
|
| 235 |
+
found: dict[str, list] = {}
|
| 236 |
+
for node in tree.body:
|
| 237 |
+
if isinstance(node, ast.Assign) and len(node.targets) == 1:
|
| 238 |
+
target = node.targets[0]
|
| 239 |
+
if isinstance(target, ast.Name) and target.id in {"VERSIONS", "DOCUMENTS", "TOKENS"}:
|
| 240 |
+
found[target.id] = ast.literal_eval(node.value)
|
| 241 |
+
missing = {"VERSIONS", "DOCUMENTS", "TOKENS"} - found.keys()
|
| 242 |
+
if missing:
|
| 243 |
+
raise LookupError(f"{path} is missing {sorted(missing)}")
|
| 244 |
+
return dict(zip(found["VERSIONS"], zip(found["DOCUMENTS"], found["TOKENS"])))
|
| 245 |
+
|
| 246 |
+
|
| 247 |
+
def load_registry(root: Path) -> dict[str, dict]:
|
| 248 |
+
"""The SOURCES table from src/sources.py, read without importing the module.
|
| 249 |
+
|
| 250 |
+
Always read from the working tree: the registry defines what a release *is*,
|
| 251 |
+
independently of the ref the parquet stats are read from.
|
| 252 |
+
"""
|
| 253 |
+
path = root / "src" / "sources.py"
|
| 254 |
+
if not path.exists():
|
| 255 |
+
return {}
|
| 256 |
+
for node in ast.parse(path.read_text(encoding="utf-8")).body:
|
| 257 |
+
if (isinstance(node, ast.Assign) and len(node.targets) == 1
|
| 258 |
+
and getattr(node.targets[0], "id", None) == "SOURCES"):
|
| 259 |
+
return ast.literal_eval(node.value)
|
| 260 |
+
return {}
|
| 261 |
+
|
| 262 |
+
|
| 263 |
+
def load_release_map(root: Path) -> dict[str, str | None]:
|
| 264 |
+
"""Source name -> the release it entered.
|
| 265 |
+
|
| 266 |
+
Empty when the registry is absent or predates the `release` field, in which
|
| 267 |
+
case the CHANGELOG prose stays the only ground truth.
|
| 268 |
+
"""
|
| 269 |
+
sources = load_registry(root)
|
| 270 |
+
if not any("release" in cfg for cfg in sources.values()):
|
| 271 |
+
return {}
|
| 272 |
+
return {name: cfg.get("release") for name, cfg in sources.items()}
|
| 273 |
+
|
| 274 |
+
|
| 275 |
+
def released_at(releases: dict[str, str | None], version: str) -> set[str]:
|
| 276 |
+
"""Sources belonging to `version`, i.e. admitted at or before it."""
|
| 277 |
+
ceiling = _version_key(version)
|
| 278 |
+
return {
|
| 279 |
+
name for name, release in releases.items()
|
| 280 |
+
if release and _version_key(release) <= ceiling
|
| 281 |
+
}
|
| 282 |
+
|
| 283 |
+
|
| 284 |
+
def _version_key(version: str) -> tuple[int, ...]:
|
| 285 |
+
return tuple(int(part) for part in version.lstrip("v").split("."))
|
| 286 |
+
|
| 287 |
+
|
| 288 |
+
def _changelog_sections(root: Path) -> dict[str, str]:
|
| 289 |
+
text = (root / "CHANGELOG.md").read_text(encoding="utf-8")
|
| 290 |
+
parts = re.split(r"^## (v[\d.]+[\w.-]*)", text, flags=re.M)
|
| 291 |
+
return dict(zip(parts[1::2], parts[2::2]))
|
| 292 |
+
|
| 293 |
+
|
| 294 |
+
def load_declared_sources(
|
| 295 |
+
root: Path, base_version: str | None = None
|
| 296 |
+
) -> tuple[dict[str, int | None], dict[str, list[str]]]:
|
| 297 |
+
"""Version -> declared source count (None when unrecorded), and sources added."""
|
| 298 |
+
sections = _changelog_sections(root)
|
| 299 |
+
counts: dict[str, int | None] = {}
|
| 300 |
+
added: dict[str, list[str]] = {}
|
| 301 |
+
for version, body in sections.items():
|
| 302 |
+
match = re.search(r"\b(\d+)[-\s]sources?\b", body)
|
| 303 |
+
counts[version] = int(match.group(1)) if match else None
|
| 304 |
+
added[version] = re.findall(
|
| 305 |
+
r"Added (?:the )?(?:community[- ]contributed |community contribution |reviewed )?`([\w.]+)`",
|
| 306 |
+
body,
|
| 307 |
+
)
|
| 308 |
+
# The first release predates the CHANGELOG; the README release table records it.
|
| 309 |
+
if base_version and counts.get(base_version) is None:
|
| 310 |
+
readme = (root / "README.md").read_text(encoding="utf-8")
|
| 311 |
+
base = re.search(
|
| 312 |
+
rf"`{re.escape(base_version)}`.*?(\d+) open/official sources", readme, re.S
|
| 313 |
+
)
|
| 314 |
+
if base:
|
| 315 |
+
counts[base_version] = int(base.group(1))
|
| 316 |
+
added.setdefault(base_version, [])
|
| 317 |
+
return counts, added
|
| 318 |
+
|
| 319 |
+
|
| 320 |
+
def load_stats(root: Path, ref: str) -> dict[str, dict]:
|
| 321 |
+
"""Source name -> parquet stats, read from a git ref or the working tree."""
|
| 322 |
+
stats: dict[str, dict] = {}
|
| 323 |
+
if ref == "worktree":
|
| 324 |
+
paths = sorted(root.glob("data/*/*.stats.json"))
|
| 325 |
+
for path in paths:
|
| 326 |
+
stats[path.name.removesuffix(".stats.json")] = json.loads(
|
| 327 |
+
path.read_text(encoding="utf-8")
|
| 328 |
+
)
|
| 329 |
+
return stats
|
| 330 |
+
|
| 331 |
+
listing = subprocess.run(
|
| 332 |
+
["git", "-C", str(root), "ls-tree", "-r", "--name-only", ref, "--", "data/"],
|
| 333 |
+
capture_output=True, text=True, check=True,
|
| 334 |
+
).stdout.splitlines()
|
| 335 |
+
for rel in sorted(p for p in listing if p.endswith(".stats.json")):
|
| 336 |
+
blob = subprocess.run(
|
| 337 |
+
["git", "-C", str(root), "show", f"{ref}:{rel}"],
|
| 338 |
+
capture_output=True, text=True, check=True,
|
| 339 |
+
).stdout
|
| 340 |
+
stats[Path(rel).name.removesuffix(".stats.json")] = json.loads(blob)
|
| 341 |
+
return stats
|
| 342 |
+
|
| 343 |
+
|
| 344 |
+
# --------------------------------------------------------------------------
|
| 345 |
+
# Checks
|
| 346 |
+
# --------------------------------------------------------------------------
|
| 347 |
+
|
| 348 |
+
def check_trajectory(report: Report, tex: str, root: Path) -> None:
|
| 349 |
+
report.section("Check 1 - release trajectory v0.2.0 -> v0.2.5 (Table 2 + hero figure)")
|
| 350 |
+
|
| 351 |
+
table = parse_trajectory_table(tex)
|
| 352 |
+
hero = parse_hero(tex)
|
| 353 |
+
series = load_growth_series(root)
|
| 354 |
+
declared, added = load_declared_sources(root, base_version=sorted(series)[0])
|
| 355 |
+
releases = load_release_map(root)
|
| 356 |
+
if releases:
|
| 357 |
+
for version in sorted(series):
|
| 358 |
+
counted = len(released_at(releases, version))
|
| 359 |
+
if declared.get(version) is None:
|
| 360 |
+
declared[version] = counted
|
| 361 |
+
elif declared[version] != counted:
|
| 362 |
+
report.add(FAIL, f"{version} registry vs CHANGELOG",
|
| 363 |
+
f"src/sources.py counts {counted}, CHANGELOG says {declared[version]}")
|
| 364 |
+
|
| 365 |
+
missing = sorted(set(series) - set(table))
|
| 366 |
+
extra = sorted(set(table) - set(series))
|
| 367 |
+
if missing or extra:
|
| 368 |
+
report.add(FAIL, "table covers every release",
|
| 369 |
+
f"missing {missing}, unexpected {extra}")
|
| 370 |
+
else:
|
| 371 |
+
report.add(OK, "table covers every release", f"{len(table)} rows")
|
| 372 |
+
|
| 373 |
+
for version in sorted(series):
|
| 374 |
+
docs, tokens = series[version]
|
| 375 |
+
actual_b = tokens / 1e9
|
| 376 |
+
claimed_b = table.get(version, (None, None))[0]
|
| 377 |
+
report.num(f"{version} tokens (B, Table 2)", claimed_b, round(actual_b, 6), TOL_BILLIONS)
|
| 378 |
+
report.num(f"{version} tokens (B, hero figure)", hero.get(version), round(actual_b, 6), TOL_BILLIONS)
|
| 379 |
+
|
| 380 |
+
# Source counts. "---" in the table means "unchanged from the previous release".
|
| 381 |
+
previous: int | None = None
|
| 382 |
+
for version in sorted(series):
|
| 383 |
+
cell = table.get(version, (None, "?"))[1]
|
| 384 |
+
if cell == "---":
|
| 385 |
+
claimed = previous
|
| 386 |
+
shown = f"--- (carries {previous})"
|
| 387 |
+
else:
|
| 388 |
+
digits = re.search(r"\d+", cell or "")
|
| 389 |
+
claimed = int(digits.group()) if digits else None
|
| 390 |
+
shown = str(claimed)
|
| 391 |
+
previous = claimed
|
| 392 |
+
|
| 393 |
+
truth = declared.get(version)
|
| 394 |
+
label = f"{version} source count"
|
| 395 |
+
if truth is None:
|
| 396 |
+
report.add(WARN, label, f"paper {shown}; no count recorded in CHANGELOG/README")
|
| 397 |
+
elif claimed == truth:
|
| 398 |
+
report.add(OK, label, f"{shown} == {truth}")
|
| 399 |
+
else:
|
| 400 |
+
report.add(FAIL, label, f"paper {shown} != repo {truth}")
|
| 401 |
+
|
| 402 |
+
if releases:
|
| 403 |
+
return # the registry already resolves every release exactly
|
| 404 |
+
|
| 405 |
+
# Back-fill unrecorded releases from the next release's "Added `source`" bullets.
|
| 406 |
+
versions = sorted(series)
|
| 407 |
+
for index, version in enumerate(versions[:-1]):
|
| 408 |
+
if declared.get(version) is not None:
|
| 409 |
+
continue
|
| 410 |
+
following = versions[index + 1]
|
| 411 |
+
if declared.get(following) is None:
|
| 412 |
+
continue
|
| 413 |
+
inferred = declared[following] - len(added.get(following, []))
|
| 414 |
+
report.add(
|
| 415 |
+
WARN,
|
| 416 |
+
f"{version} source count (derived)",
|
| 417 |
+
f"{following} declares {declared[following]} and adds "
|
| 418 |
+
f"{len(added.get(following, []))} -> {version} = {inferred}",
|
| 419 |
+
)
|
| 420 |
+
|
| 421 |
+
|
| 422 |
+
def check_composition(report: Report, tex: str, root: Path, ref: str) -> None:
|
| 423 |
+
report.section(f"Check 2 - corpus composition and totals (Table 3, ref={ref})")
|
| 424 |
+
|
| 425 |
+
table, table_total = parse_source_table(tex)
|
| 426 |
+
current = sorted(load_growth_series(root))[-1]
|
| 427 |
+
claims = parse_inline_claims(tex, current=current)
|
| 428 |
+
stats = load_stats(root, ref)
|
| 429 |
+
releases = load_release_map(root)
|
| 430 |
+
in_release = released_at(releases, current) if releases else set(stats)
|
| 431 |
+
|
| 432 |
+
absent = sorted(set(table) - set(stats))
|
| 433 |
+
unlisted = sorted((set(stats) & in_release) - set(table))
|
| 434 |
+
unreleased = sorted(set(stats) - in_release)
|
| 435 |
+
listed_but_unreleased = sorted(set(table) & set(unreleased))
|
| 436 |
+
if absent:
|
| 437 |
+
report.add(FAIL, "every listed source exists at ref", f"missing from repo: {absent}")
|
| 438 |
+
else:
|
| 439 |
+
report.add(OK, "every listed source exists at ref", f"{len(table)} sources")
|
| 440 |
+
if unlisted:
|
| 441 |
+
report.add(FAIL, f"every {current} source is in Table 3",
|
| 442 |
+
f"released but missing from Table 3: {unlisted}")
|
| 443 |
+
else:
|
| 444 |
+
report.add(OK, f"every {current} source is in Table 3", "")
|
| 445 |
+
if listed_but_unreleased:
|
| 446 |
+
report.add(FAIL, "Table 3 lists only released sources",
|
| 447 |
+
f"not in any release: {listed_but_unreleased}")
|
| 448 |
+
elif unreleased:
|
| 449 |
+
report.add(OK, "Table 3 lists only released sources",
|
| 450 |
+
f"correctly excludes {unreleased}")
|
| 451 |
+
|
| 452 |
+
listed = [name for name in table if name in stats]
|
| 453 |
+
for name in sorted(listed, key=lambda n: -stats[n]["tokens"]):
|
| 454 |
+
report.num(f"{name} tokens (M)", table[name], round(stats[name]["tokens"] / 1e6, 6), TOL_MILLIONS)
|
| 455 |
+
|
| 456 |
+
sum_tokens = sum(stats[name]["tokens"] for name in listed)
|
| 457 |
+
sum_docs = sum(stats[name]["kept"] for name in listed)
|
| 458 |
+
|
| 459 |
+
report.num("Table 3 rows sum to its Total row (M)",
|
| 460 |
+
round(sum(table[n] for n in listed), 6),
|
| 461 |
+
round(sum_tokens / 1e6, 6), TOL_MILLIONS * len(listed))
|
| 462 |
+
report.num("Table 3 Total row (M)", table_total, round(sum_tokens / 1e6, 6), 1.0)
|
| 463 |
+
|
| 464 |
+
for value in sorted(claims["documents"]):
|
| 465 |
+
report.num(f"prose claim: {value} documents", value, sum_docs)
|
| 466 |
+
for value in sorted(claims["tokens_billions"]):
|
| 467 |
+
report.num(f"prose claim: {value}B tokens", value, round(sum_tokens / 1e9, 6), TOL_BILLIONS)
|
| 468 |
+
for value in sorted(claims["sources"]):
|
| 469 |
+
report.num(f"prose claim: {value} sources", value, len(listed))
|
| 470 |
+
|
| 471 |
+
if unreleased:
|
| 472 |
+
report.add(
|
| 473 |
+
WARN, "totals over ALL sources at ref",
|
| 474 |
+
f"{sum(s['kept'] for s in stats.values())} docs / "
|
| 475 |
+
f"{sum(s['tokens'] for s in stats.values()) / 1e9:.3f}B tokens "
|
| 476 |
+
f"across {len(stats)} sources",
|
| 477 |
+
)
|
| 478 |
+
|
| 479 |
+
# Anchor: the trajectory series' last point must equal the composition sum.
|
| 480 |
+
series = load_growth_series(root)
|
| 481 |
+
last = sorted(series)[-1]
|
| 482 |
+
report.num(f"{last} tokens in growth series == Table 3 sum",
|
| 483 |
+
series[last][1], sum_tokens)
|
| 484 |
+
report.num(f"{last} documents in growth series == Table 3 sum",
|
| 485 |
+
series[last][0], sum_docs)
|
| 486 |
+
|
| 487 |
+
|
| 488 |
+
def norm_license(text: str) -> tuple[str, str | None]:
|
| 489 |
+
"""Canonical (family, version) so paper, registry and datasheet can be compared.
|
| 490 |
+
|
| 491 |
+
Public domain and CC0 are one family: SPDX spells "official documents are
|
| 492 |
+
outside copyright" as CC0-1.0, and the paper prints both as PD. A version of
|
| 493 |
+
None means the text states no version, and compares equal to any version --
|
| 494 |
+
the paper abbreviates nkjp1m's CC-BY-4.0 to "CC-BY" on purpose.
|
| 495 |
+
"""
|
| 496 |
+
t = re.sub(r"[`_]", " ", text).upper()
|
| 497 |
+
if re.search(r"PER-RECORD|PER-DOCUMENT|MIXED", t):
|
| 498 |
+
return ("MIXED", None)
|
| 499 |
+
if re.search(r"CC0|PUBLIC[- ]DOMAIN|\bPD\b", t):
|
| 500 |
+
return ("PD/CC0", None)
|
| 501 |
+
match = re.search(r"CC[- ]BY(?:[- ](SA))?(?:[- ](\d+\.\d+))?", t)
|
| 502 |
+
if match:
|
| 503 |
+
return ("CC-BY-SA" if match.group(1) else "CC-BY", match.group(2))
|
| 504 |
+
return (t.strip(), None)
|
| 505 |
+
|
| 506 |
+
|
| 507 |
+
def compatible(a: tuple[str, str | None], b: tuple[str, str | None]) -> bool:
|
| 508 |
+
"""Same family, and same version unless one side states none."""
|
| 509 |
+
return a[0] == b[0] and (a[1] is None or b[1] is None or a[1] == b[1])
|
| 510 |
+
|
| 511 |
+
|
| 512 |
+
def load_datasheet_license(root: Path, name: str) -> str | None:
|
| 513 |
+
"""The datasheet's declared License line (working tree, not the pinned ref)."""
|
| 514 |
+
path = root / "data" / name / f"{name}.md"
|
| 515 |
+
if not path.exists():
|
| 516 |
+
return None
|
| 517 |
+
match = re.search(r"\*\*License:\*\*\s*(.+)", path.read_text(encoding="utf-8"))
|
| 518 |
+
return match.group(1).strip() if match else None
|
| 519 |
+
|
| 520 |
+
|
| 521 |
+
def check_families(report: Report, tex: str, root: Path, ref: str) -> None:
|
| 522 |
+
"""Table 5 must be a complete partition of Table 3 by license family."""
|
| 523 |
+
paper, paper_total = parse_family_table(tex)
|
| 524 |
+
registry, stats = load_registry(root), load_stats(root, ref)
|
| 525 |
+
releases = load_release_map(root)
|
| 526 |
+
if not (paper and registry and releases):
|
| 527 |
+
report.add(WARN, "license family table", "no family table or no registry to partition")
|
| 528 |
+
return
|
| 529 |
+
|
| 530 |
+
repo: dict[str, int] = {}
|
| 531 |
+
unclassified = []
|
| 532 |
+
for name in released_at(releases, sorted(load_growth_series(root))[-1]):
|
| 533 |
+
cfg = registry[name]
|
| 534 |
+
# The row prints the content's own status; SPDX carries the CC0 wording.
|
| 535 |
+
family = license_family(str(cfg.get("license", ""))) or license_family(str(cfg.get("license_spdx") or ""))
|
| 536 |
+
if family is None:
|
| 537 |
+
unclassified.append(name)
|
| 538 |
+
else:
|
| 539 |
+
repo[family] = repo.get(family, 0) + stats[name]["tokens"]
|
| 540 |
+
|
| 541 |
+
if unclassified:
|
| 542 |
+
report.add(FAIL, "every source has a license family", f"unclassifiable: {sorted(unclassified)}")
|
| 543 |
+
if set(paper) != set(repo):
|
| 544 |
+
report.add(FAIL, "Table 5 families match the corpus",
|
| 545 |
+
f"paper {sorted(paper)} != repo {sorted(repo)}")
|
| 546 |
+
else:
|
| 547 |
+
report.add(OK, "Table 5 families match the corpus", f"{len(repo)} families, no source left over")
|
| 548 |
+
|
| 549 |
+
for family in sorted(repo, key=lambda f: -repo[f]):
|
| 550 |
+
if family in paper:
|
| 551 |
+
report.num(f"{family} tokens (B)", paper[family], round(repo[family] / 1e9, 6), TOL_BILLIONS)
|
| 552 |
+
report.num("Table 5 Total row (B)", paper_total, round(sum(repo.values()) / 1e9, 6), TOL_BILLIONS)
|
| 553 |
+
|
| 554 |
+
|
| 555 |
+
def check_licenses(report: Report, tex: str, root: Path) -> None:
|
| 556 |
+
report.section("Check 3 - licenses, footnotes and license families")
|
| 557 |
+
|
| 558 |
+
rows = parse_license_rows(tex)
|
| 559 |
+
registry = load_registry(root)
|
| 560 |
+
if not registry:
|
| 561 |
+
report.add(WARN, "license audit", "src/sources.py carries no registry to compare against")
|
| 562 |
+
return
|
| 563 |
+
|
| 564 |
+
for name in sorted(rows, key=lambda n: n.lower()):
|
| 565 |
+
paper, markers = rows[name]
|
| 566 |
+
cfg = registry.get(name)
|
| 567 |
+
if cfg is None:
|
| 568 |
+
report.add(FAIL, f"{name} in registry", "listed in Table 3 but absent from src/sources.py")
|
| 569 |
+
continue
|
| 570 |
+
|
| 571 |
+
declared = norm_license(str(cfg.get("license", "")))
|
| 572 |
+
spdx = norm_license(str(cfg.get("license_spdx") or cfg.get("license", "")))
|
| 573 |
+
printed = norm_license(paper)
|
| 574 |
+
|
| 575 |
+
if compatible(printed, declared) or compatible(printed, spdx):
|
| 576 |
+
report.add(OK, f"{name} license", f"{paper} == {cfg.get('license')}")
|
| 577 |
+
else:
|
| 578 |
+
report.add(FAIL, f"{name} license",
|
| 579 |
+
f"paper {paper!r} != registry {cfg.get('license')!r} / {cfg.get('license_spdx')!r}")
|
| 580 |
+
|
| 581 |
+
sheet = load_datasheet_license(root, name)
|
| 582 |
+
if sheet is None:
|
| 583 |
+
report.add(WARN, f"{name} datasheet license", "no **License:** line in the datasheet")
|
| 584 |
+
elif not compatible(norm_license(sheet), declared):
|
| 585 |
+
report.add(FAIL, f"{name} datasheet license",
|
| 586 |
+
f"datasheet {sheet!r} != registry {cfg.get('license')!r}")
|
| 587 |
+
|
| 588 |
+
# A source whose content license differs from its redistribution license
|
| 589 |
+
# states only half the story in a one-line cell; the footnote carries the rest.
|
| 590 |
+
if not compatible(declared, spdx) and not markers:
|
| 591 |
+
report.add(FAIL, f"{name} dual license is footnoted",
|
| 592 |
+
f"content {cfg.get('license')!r} but redistributed as "
|
| 593 |
+
f"{cfg.get('license_spdx')!r}; the row prints only {paper!r}")
|
| 594 |
+
|
| 595 |
+
used = {m for _, markers in rows.values() for m in markers}
|
| 596 |
+
defined = parse_footnote_defs(tex)
|
| 597 |
+
if used - defined:
|
| 598 |
+
report.add(FAIL, "every footnote marker is defined", f"used but not explained: {sorted(used - defined)}")
|
| 599 |
+
elif defined - used:
|
| 600 |
+
report.add(FAIL, "every footnote is used", f"explained but on no row: {sorted(defined - used)}")
|
| 601 |
+
else:
|
| 602 |
+
report.add(OK, "footnote markers", f"{len(used)} defined and used: {sorted(used)}")
|
| 603 |
+
|
| 604 |
+
total = parse_total_license(tex)
|
| 605 |
+
strongest = [n for n, (lic, _) in rows.items() if norm_license(lic)[0] == "CC-BY-SA"]
|
| 606 |
+
if total is None:
|
| 607 |
+
report.add(WARN, "Total row license", "no license cell in the Total row")
|
| 608 |
+
elif strongest and norm_license(total)[0] != "CC-BY-SA":
|
| 609 |
+
report.add(FAIL, "Total row license", f"{total} does not inherit the share-alike of {sorted(strongest)}")
|
| 610 |
+
else:
|
| 611 |
+
report.add(OK, "Total row license", f"{total} inherits share-alike from {len(strongest)} sources")
|
| 612 |
+
|
| 613 |
+
|
| 614 |
+
def main(argv: list[str] | None = None) -> int:
|
| 615 |
+
parser = argparse.ArgumentParser(description=__doc__,
|
| 616 |
+
formatter_class=argparse.RawDescriptionHelpFormatter)
|
| 617 |
+
parser.add_argument("--tex", default=str(ROOT / "file.tex"), help="paper source")
|
| 618 |
+
parser.add_argument("--root", default=str(ROOT), help="repository root")
|
| 619 |
+
parser.add_argument("--ref", default=None,
|
| 620 |
+
help="git ref for data/*.stats.json, or 'worktree' "
|
| 621 |
+
"(default: the commit the .tex pins)")
|
| 622 |
+
args = parser.parse_args(argv)
|
| 623 |
+
|
| 624 |
+
root = Path(args.root).resolve()
|
| 625 |
+
tex = Path(args.tex).read_text(encoding="utf-8")
|
| 626 |
+
ref = args.ref or parse_inline_claims(tex)["pin"] or "HEAD"
|
| 627 |
+
|
| 628 |
+
report = Report()
|
| 629 |
+
check_trajectory(report, tex, root)
|
| 630 |
+
check_composition(report, tex, root, ref)
|
| 631 |
+
check_licenses(report, tex, root)
|
| 632 |
+
check_families(report, tex, root, ref)
|
| 633 |
+
return report.render()
|
| 634 |
+
|
| 635 |
+
|
| 636 |
+
if __name__ == "__main__":
|
| 637 |
+
sys.exit(main())
|