Maggio33 ppuzio commited on
Commit
77e5321
·
1 Parent(s): a916ede

Make release totals registry-driven (#29)

Browse files

- fix: make release totals registry-driven, not "whatever is on main" (4f879612263127710c44a0c376b697088f2ca146)
- docs: how to push and open PRs on the Hub (10da95c3d913e08491106f1b6628b59e15105738)
- feat: audit source licenses and table footnotes (4fbe51379b807370e4268e0a7b2809ffc5ac43da)
- feat: verify the license-family table is a complete partition (d0b620e22be4c888ae3851602ba012961eb61245)


Co-authored-by: Paweł Puzio <ppuzio@users.noreply.huggingface.co>

AGENTS.md CHANGED
@@ -13,8 +13,9 @@ committed docs on its own:
13
 
14
  - It iterates `SOURCES` in `src/sources.py`; any source in `data/` but missing
15
  from `SOURCES` is silently dropped from the table and totals.
16
- - It stamps every regenerated datasheet's "Added" with the single global `ADDED`
17
- constant, overwriting real per-source add dates.
 
18
  - It cannot reproduce hand-written README narrative (audit notes, policy notes,
19
  the phrase-frequency section).
20
 
@@ -44,6 +45,11 @@ answer was yes or no.
44
  `speakleash_key`/`file_key`. If the datasheet is hand-authored (rich provenance
45
  or a fixed add date), add `"custom_datasheet": True` so make_docs leaves it
46
  alone.
 
 
 
 
 
47
  4. Add a contract test: `src/test_<key>_contract.py` (canonical schema, non-empty
48
  text, positive token counts, uniform source/license, stats-file consistency).
49
  Run `python3 -m pytest src/` before committing.
@@ -142,6 +148,47 @@ Two rules about all of the above:
142
  4. Prepend the release notes to `CHANGELOG.md` under a new `## vX.Y.Z (date)`
143
  heading. History is append-only — never rewrite earlier releases.
144
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
145
  ## Do not commit
146
 
147
  `*.log` (fetch/build logs), `.DS_Store`, `src/__pycache__/`. Only `logs/` is in
@@ -150,8 +197,7 @@ before staging. `*.parquet` is LFS-tracked; commit the pointer, not the blob.
150
 
151
  ## Known drift / cleanup opportunities
152
 
153
- - `biblioteka_nauki`, `europeana`, `parlamint_pl` are in `SOURCES` but have no
154
- `data/` shard (make_docs skips them with `! no stats`). Intentional placeholders
155
- or stale — confirm before relying on `build_dynaword.py --all`.
156
- - The global `ADDED` datasheet stamp is a footgun: give datasheets a per-source
157
- add date (or `custom_datasheet`) before making make_docs regenerate them.
 
13
 
14
  - It iterates `SOURCES` in `src/sources.py`; any source in `data/` but missing
15
  from `SOURCES` is silently dropped from the table and totals.
16
+ - It stamps a regenerated datasheet's "Added" with the source's `added` field,
17
+ falling back to the global `ADDED` constant when the entry has none — so a new
18
+ entry without `added` still gets a wrong date.
19
  - It cannot reproduce hand-written README narrative (audit notes, policy notes,
20
  the phrase-frequency section).
21
 
 
45
  `speakleash_key`/`file_key`. If the datasheet is hand-authored (rich provenance
46
  or a fixed add date), add `"custom_datasheet": True` so make_docs leaves it
47
  alone.
48
+
49
+ Also set `added` (the real add date) and `release`. `release` is what admits a
50
+ source to a release's totals: `None` means "on main, not in any release yet",
51
+ which is the correct value for a new source until the release that ships it is
52
+ cut. Nothing is counted just because it has a `stats.json`.
53
  4. Add a contract test: `src/test_<key>_contract.py` (canonical schema, non-empty
54
  text, positive token counts, uniform source/license, stats-file consistency).
55
  Run `python3 -m pytest src/` before committing.
 
148
  4. Prepend the release notes to `CHANGELOG.md` under a new `## vX.Y.Z (date)`
149
  heading. History is append-only — never rewrite earlier releases.
150
 
151
+ ## Push to Hugging Face
152
+
153
+ `origin` is the Hub itself
154
+ (`https://huggingface.co/datasets/SlayerLab/polish-dynaword`) — there is no GitHub
155
+ remote. `git push` authenticates through the git credential helper (macOS:
156
+ `osxkeychain`), which is separate from `hf auth login`: pushes can work fine while
157
+ `hf auth whoami` still reports "Not logged in". The Python API and `hf` CLI need
158
+ the login (or `HF_TOKEN`).
159
+
160
+ Hub pull requests use no forks and no named branches. A PR *is* the ref
161
+ `refs/pr/N` and the Hub assigns `N`, so you cannot open one by pushing — the ref
162
+ has to exist first. (Plain branches can be pushed, but a branch is not a PR and
163
+ nobody reviews it.)
164
+
165
+ Open a PR for work already committed locally:
166
+
167
+ ```bash
168
+ # 1. create the empty PR (needs the write token)
169
+ python3 -c "
170
+ from huggingface_hub import HfApi
171
+ pr = HfApi().create_pull_request(
172
+ 'SlayerLab/polish-dynaword', repo_type='dataset',
173
+ title='<title>', description='<what changed and why>')
174
+ print(pr.num, pr.url)"
175
+
176
+ # 2. push your local branch onto that ref
177
+ git push origin <local-branch>:refs/pr/<N>
178
+ ```
179
+
180
+ Opened this way the PR starts in **draft** — publish it from the web UI.
181
+
182
+ Pick up an existing PR (42 here):
183
+
184
+ ```bash
185
+ git fetch origin refs/pr/42:pr/42 && git checkout pr/42
186
+ git push origin pr/42:refs/pr/42
187
+ ```
188
+
189
+ `repo_type='dataset'` is required on every `huggingface_hub` call — this is a
190
+ dataset repo, not a model. Push straight to `main` only when cutting a release.
191
+
192
  ## Do not commit
193
 
194
  `*.log` (fetch/build logs), `.DS_Store`, `src/__pycache__/`. Only `logs/` is in
 
197
 
198
  ## Known drift / cleanup opportunities
199
 
200
+ - `europeana` is in `SOURCES` but has no `data/` shard anywhere, and
201
+ `parlamint_pl` has one only in some working trees — it is not committed.
202
+ make_docs skips both with `! no stats`. Intentional placeholders or stale —
203
+ confirm before relying on `build_dynaword.py --all`.
 
CHANGELOG.md CHANGED
@@ -1,5 +1,14 @@
1
  # Changelog
2
 
 
 
 
 
 
 
 
 
 
3
  ## v0.2.5 (2026-08-14)
4
 
5
  - Added the community contribution `samorzad_gov_pl` by Dawid Majewski
@@ -43,8 +52,9 @@
43
  - Expanded contributor credit for Arkadiusz Słota (`Maggio33`) for the WDS 10-5
44
  expansion, PII-scrub and dedup work.
45
  - Updated the complete-corpus release totals from per-source parquet stats:
46
- 4,246,429 documents and 9,597,908,067 cl100k-proxy tokens (also reflects the
47
- separately-merged `biblioteka_nauki` source now present on main).
 
48
 
49
  ## v0.2.3 (2026-07-26)
50
 
 
1
  # Changelog
2
 
3
+ ## Unreleased (on main)
4
+
5
+ - `wiktionary_examples`, `sejm_api` and `parlamint_pl` are built, documented and
6
+ present on the main branch but have never been admitted to a stable release.
7
+ They are excluded from every release total until a release entry admits them.
8
+ The registry records this as `release: None` in `src/sources.py`; release
9
+ totals count only sources whose `release` is set and not newer than the
10
+ version being documented.
11
+
12
  ## v0.2.5 (2026-08-14)
13
 
14
  - Added the community contribution `samorzad_gov_pl` by Dawid Majewski
 
52
  - Expanded contributor credit for Arkadiusz Słota (`Maggio33`) for the WDS 10-5
53
  expansion, PII-scrub and dedup work.
54
  - Updated the complete-corpus release totals from per-source parquet stats:
55
+ 4,246,429 documents and 9,597,908,067 cl100k-proxy tokens across 17 sources
56
+ (also reflects the separately-merged `biblioteka_nauki` source now present on
57
+ main).
58
 
59
  ## v0.2.3 (2026-07-26)
60
 
README.md CHANGED
@@ -182,7 +182,16 @@ files is not listed as a data contribution.
182
  | [wikinews](data/wikinews/wikinews.md) | Polish Wikinews | `CC-BY-2.5` | 24,386 | 12.1M |
183
  | [global_voices](data/global_voices/global_voices.md) | Global Voices Polish | `CC-BY-3.0` | 2,040 | 3.7M |
184
  | [nkjp1m](data/nkjp1m/nkjp1m.md) | The manually annotated 1-million word subcorpus of the National Corpus of Polish | `CC-BY` | 18,381 | 2.6M |
185
- | **total** | | | **4,246,429** | **9,597.9M** |
 
 
 
 
 
 
 
 
 
186
 
187
  ## Method
188
  Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
 
182
  | [wikinews](data/wikinews/wikinews.md) | Polish Wikinews | `CC-BY-2.5` | 24,386 | 12.1M |
183
  | [global_voices](data/global_voices/global_voices.md) | Global Voices Polish | `CC-BY-3.0` | 2,040 | 3.7M |
184
  | [nkjp1m](data/nkjp1m/nkjp1m.md) | The manually annotated 1-million word subcorpus of the National Corpus of Polish | `CC-BY` | 18,381 | 2.6M |
185
+ | **total** | | | **4,319,200** | **9,639.1M** |
186
+
187
+ ## Sources on main, not yet released
188
+ Built and documented, but **not** part of v0.2.5 and excluded from every total above. They join a release when a CHANGELOG entry admits them.
189
+
190
+ | source | description | license | documents | tokens |
191
+ |---|---|---|---:|---:|
192
+ | [parlamint_pl](data/parlamint_pl/parlamint_pl.md) | ParlaMint-PL (parliamentary debates, 2015-2022) | `CC-BY-4.0` | 686 | 95.4M |
193
+ | [sejm_api](data/sejm_api/sejm_api.md) | Sejm API parliamentary speeches (2023 onward) | `public-domain (official documents)` | 38,812 | 36.3M |
194
+ | [wiktionary_examples](data/wiktionary_examples/wiktionary_examples.md) | Polish Wiktionary usage examples | `CC-BY-SA-3.0` | 8,639 | 1.1M |
195
 
196
  ## Method
197
  Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
src/make_docs.py CHANGED
@@ -2,6 +2,9 @@
2
  """Generate Dynaword documentation: per-source datasheets + README + CHANGELOG + LICENSE.
3
 
4
  Reads sources.py + data/<source>/<source>.stats.json (written by build_dynaword.py).
 
 
 
5
  Implements the "Documented" principle (datasheets, Gebru et al. 2021) and the
6
  aggregate README table (paper 2508.02271).
7
  """
@@ -49,12 +52,31 @@ from the next version (see retroactive-removal policy below).
49
  """
50
 
51
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
52
  def load_stats(name):
53
  f = ROOT / "data" / name / f"{name}.stats.json"
54
  return json.loads(f.read_text()) if f.exists() else None
55
 
56
 
57
  def datasheet(name, cfg, st):
 
58
  license_rows = ""
59
  licenses = st.get("licenses") or {}
60
  if licenses:
@@ -84,7 +106,7 @@ def datasheet(name, cfg, st):
84
  - **Language:** Polish (pl)
85
  - **License:** `{cfg['license']}`
86
  - **Created (range):** {cfg['created']}
87
- - **Added:** {ADDED}
88
 
89
  ## Licensing — traceable basis
90
  {cfg['traceable']}
@@ -115,7 +137,8 @@ Llama-3 count is computed at release.
115
 
116
 
117
  def main():
118
- rows, tot_doc, tot_tok, tot_chr = [], 0, 0, 0
 
119
  for name, cfg in SOURCES.items():
120
  st = load_stats(name)
121
  if not st:
@@ -123,16 +146,36 @@ def main():
123
  # Sources with a hand-authored datasheet (rich provenance, per-source add
124
  # date) opt out of template regeneration to avoid clobbering it.
125
  if not cfg.get("custom_datasheet"):
126
- (ROOT / "data" / name / f"{name}.md").write_text(datasheet(name, cfg, st))
 
 
 
 
 
127
  rows.append((name, cfg, st))
128
  tot_doc += st["kept"]; tot_tok += st["tokens"]; tot_chr += st["chars"]
129
  rows.sort(key=lambda r: -r[2]["tokens"])
 
130
 
131
  tbl = "\n".join(
132
  f"| [{n}](data/{n}/{n}.md) | {c['pretty']} | `{c['license']}` | "
133
  f"{s['kept']:,} | {s['tokens']/1e6:,.1f}M |"
134
  for n, c, s in rows)
135
  excl = "\n".join(f"| `{k}` | {v} |" for k, v in EXCLUDED.items())
 
 
 
 
 
 
 
 
 
 
 
 
 
 
136
  phrase_frequency = ""
137
  phrase_path = ROOT / "artifacts" / "pattern_frequency_hf_snippet.md"
138
  if phrase_path.exists():
@@ -143,7 +186,7 @@ def main():
143
  "\n## Results\n\n"
144
  "### Corpus phrase frequency (normalized by tokens)\n\n"
145
  "Raw counts and token-normalized shares are regenerated from the "
146
- "current parquet files with `src/pattern_frequency_report.py`.\n\n"
147
  f"{snippet}"
148
  )
149
 
@@ -206,11 +249,11 @@ family (Enevoldsen et al., [arXiv:2508.02271](https://arxiv.org/abs/2508.02271))
206
 
207
  #### v0.2.5 — current stable
208
 
209
- Added `samorzad_gov_pl`, a directly collected corpus of articles from 246
210
- Polish local public institutions and one central platform tenant. The
211
- contribution preserves the publishing institution as `author`, discovered and
212
- fetched article URLs in an attribution sidecar and the platform's CC-BY-SA-4.0
213
- basis.
214
 
215
  #### v0.2.4 — previous stable
216
 
@@ -315,7 +358,7 @@ files is not listed as a data contribution.
315
  |---|---|---|---:|---:|
316
  {tbl}
317
  | **total** | | | **{tot_doc:,}** | **{tot_tok/1e6:,.1f}M** |
318
-
319
  ## Method
320
  Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
321
  data. Gates are intentionally minimal (drop short docs, non-Polish, exact duplicates,
@@ -389,7 +432,7 @@ python3 src/make_docs.py
389
  ```
390
  {phrase_frequency}
391
  """
392
- (ROOT / "README.md").write_text(readme)
393
 
394
  # CHANGELOG is append-only release history. Documentation regeneration must
395
  # never erase earlier releases or hand-reviewed legal/release notes.
@@ -403,15 +446,19 @@ python3 src/make_docs.py
403
  f"{tot_doc:,} docs, {tot_tok:,} tokens (tiktoken cl100k proxy).\n\n"
404
  )
405
  changelog = changelog.replace("# Changelog\n", f"# Changelog\n\n{entry}", 1)
406
- changelog_path.write_text(changelog)
407
 
408
- (ROOT / "LICENSE").write_text(
 
409
  "Polish DynaWord is released under Creative Commons Attribution-ShareAlike 4.0\n"
410
  "International (CC-BY-SA-4.0): https://creativecommons.org/licenses/by-sa/4.0/\n\n"
411
  "Per-source upstream licenses and attribution are documented in each\n"
412
  "data/<source>/<source>.md datasheet.\n")
413
 
414
- print(f"docs written: README + CHANGELOG + LICENSE + {len(rows)} datasheets")
 
 
 
415
  print(f"TOTAL {tot_doc:,} docs | {tot_tok/1e9:.2f}B tok | {tot_chr/1e9:.1f}B chars")
416
 
417
 
 
2
  """Generate Dynaword documentation: per-source datasheets + README + CHANGELOG + LICENSE.
3
 
4
  Reads sources.py + data/<source>/<source>.stats.json (written by build_dynaword.py).
5
+ Only sources whose registry `release` is set and <= VERSION count toward release
6
+ totals; sources on main awaiting a release are listed separately and excluded, so
7
+ merging a new source cannot silently redefine a published total.
8
  Implements the "Documented" principle (datasheets, Gebru et al. 2021) and the
9
  aggregate README table (paper 2508.02271).
10
  """
 
52
  """
53
 
54
 
55
+ def write_text(path, text):
56
+ """Write UTF-8, preserving the file's existing line endings.
57
+
58
+ README.md and CHANGELOG.md are stored CRLF (last regenerated on Windows).
59
+ Writing LF from another OS would rewrite every line, burying the real change
60
+ in a whole-file whitespace diff.
61
+ """
62
+ text = text.replace("\r\n", "\n")
63
+ if path.exists() and b"\r\n" in path.read_bytes():
64
+ text = text.replace("\n", "\r\n")
65
+ path.write_bytes(text.encode("utf-8"))
66
+
67
+
68
+ def parse_version(text):
69
+ """'0.2.10' -> (0, 2, 10), so releases order numerically rather than lexically."""
70
+ return tuple(int(part) for part in text.split("."))
71
+
72
+
73
  def load_stats(name):
74
  f = ROOT / "data" / name / f"{name}.stats.json"
75
  return json.loads(f.read_text()) if f.exists() else None
76
 
77
 
78
  def datasheet(name, cfg, st):
79
+ added = cfg.get("added") or ADDED
80
  license_rows = ""
81
  licenses = st.get("licenses") or {}
82
  if licenses:
 
106
  - **Language:** Polish (pl)
107
  - **License:** `{cfg['license']}`
108
  - **Created (range):** {cfg['created']}
109
+ - **Added:** {added}
110
 
111
  ## Licensing — traceable basis
112
  {cfg['traceable']}
 
137
 
138
 
139
  def main():
140
+ current = parse_version(VERSION)
141
+ rows, pending, tot_doc, tot_tok, tot_chr = [], [], 0, 0, 0
142
  for name, cfg in SOURCES.items():
143
  st = load_stats(name)
144
  if not st:
 
146
  # Sources with a hand-authored datasheet (rich provenance, per-source add
147
  # date) opt out of template regeneration to avoid clobbering it.
148
  if not cfg.get("custom_datasheet"):
149
+ write_text(ROOT / "data" / name / f"{name}.md", datasheet(name, cfg, st))
150
+ release = cfg.get("release")
151
+ if release is None or parse_version(release) > current:
152
+ # Built and on main, but not admitted to this release: never counted.
153
+ pending.append((name, cfg, st))
154
+ continue
155
  rows.append((name, cfg, st))
156
  tot_doc += st["kept"]; tot_tok += st["tokens"]; tot_chr += st["chars"]
157
  rows.sort(key=lambda r: -r[2]["tokens"])
158
+ pending.sort(key=lambda r: -r[2]["tokens"])
159
 
160
  tbl = "\n".join(
161
  f"| [{n}](data/{n}/{n}.md) | {c['pretty']} | `{c['license']}` | "
162
  f"{s['kept']:,} | {s['tokens']/1e6:,.1f}M |"
163
  for n, c, s in rows)
164
  excl = "\n".join(f"| `{k}` | {v} |" for k, v in EXCLUDED.items())
165
+ unreleased = ""
166
+ if pending:
167
+ unreleased_rows = "\n".join(
168
+ f"| [{n}](data/{n}/{n}.md) | {c['pretty']} | `{c['license']}` | "
169
+ f"{s_['kept']:,} | {s_['tokens']/1e6:,.1f}M |"
170
+ for n, c, s_ in pending)
171
+ unreleased = (
172
+ "\n## Sources on main, not yet released\n"
173
+ f"Built and documented, but **not** part of v{VERSION} and excluded from "
174
+ "every total above. They join a release when a CHANGELOG entry admits "
175
+ "them.\n\n"
176
+ "| source | description | license | documents | tokens |\n"
177
+ "|---|---|---|---:|---:|\n"
178
+ f"{unreleased_rows}\n")
179
  phrase_frequency = ""
180
  phrase_path = ROOT / "artifacts" / "pattern_frequency_hf_snippet.md"
181
  if phrase_path.exists():
 
186
  "\n## Results\n\n"
187
  "### Corpus phrase frequency (normalized by tokens)\n\n"
188
  "Raw counts and token-normalized shares are regenerated from the "
189
+ "current Parquet files with `src/pattern_frequency_report.py`.\n\n"
190
  f"{snippet}"
191
  )
192
 
 
249
 
250
  #### v0.2.5 — current stable
251
 
252
+ Added `samorzad_gov_pl`: **72,771 documents / 41,154,506 cl100k-proxy
253
+ tokens** from 246 Polish local public institutions and one central platform
254
+ tenant. The contribution preserves the publishing institution as `author`,
255
+ discovered and fetched article URLs in an attribution sidecar and the verified
256
+ platform-wide CC-BY-SA-4.0 basis.
257
 
258
  #### v0.2.4 — previous stable
259
 
 
358
  |---|---|---|---:|---:|
359
  {tbl}
360
  | **total** | | | **{tot_doc:,}** | **{tot_tok/1e6:,.1f}M** |
361
+ {unreleased}
362
  ## Method
363
  Only **human-authored** text — no synthetic, machine-translated, or auto-transcribed
364
  data. Gates are intentionally minimal (drop short docs, non-Polish, exact duplicates,
 
432
  ```
433
  {phrase_frequency}
434
  """
435
+ write_text(ROOT / "README.md", readme.rstrip("\n") + "\n")
436
 
437
  # CHANGELOG is append-only release history. Documentation regeneration must
438
  # never erase earlier releases or hand-reviewed legal/release notes.
 
446
  f"{tot_doc:,} docs, {tot_tok:,} tokens (tiktoken cl100k proxy).\n\n"
447
  )
448
  changelog = changelog.replace("# Changelog\n", f"# Changelog\n\n{entry}", 1)
449
+ write_text(changelog_path, changelog)
450
 
451
+ write_text(
452
+ ROOT / "LICENSE",
453
  "Polish DynaWord is released under Creative Commons Attribution-ShareAlike 4.0\n"
454
  "International (CC-BY-SA-4.0): https://creativecommons.org/licenses/by-sa/4.0/\n\n"
455
  "Per-source upstream licenses and attribution are documented in each\n"
456
  "data/<source>/<source>.md datasheet.\n")
457
 
458
+ print(f"docs written: README + CHANGELOG + LICENSE + "
459
+ f"{len(rows) + len(pending)} datasheets")
460
+ print(f"release v{VERSION}: {len(rows)} sources"
461
+ + (f" | on main, unreleased: {len(pending)}" if pending else ""))
462
  print(f"TOTAL {tot_doc:,} docs | {tot_tok/1e9:.2f}B tok | {tot_chr/1e9:.1f}B chars")
463
 
464
 
src/sources.py CHANGED
@@ -16,6 +16,8 @@ intermediate aggregator, upstream license/attribution is preserved per source.
16
  # be rebuilt from direct upstream/export scripts rather than blind SpeakLeash pulls.
17
  SOURCES = {
18
  "wikipedia": {
 
 
19
  "speakleash_key": "plwiki",
20
  "pretty": "Polish Wikipedia",
21
  "license": "CC-BY-SA-3.0",
@@ -28,6 +30,8 @@ SOURCES = {
28
  "is_ocr": False,
29
  },
30
  "wikisource": {
 
 
31
  "speakleash_key": "plwikisource",
32
  "pretty": "Polish Wikisource",
33
  "license": "CC-BY-SA-3.0",
@@ -40,6 +44,8 @@ SOURCES = {
40
  "is_ocr": False,
41
  },
42
  "wiktionary_examples": {
 
 
43
  "file_key": "wiktionary_examples",
44
  "pretty": "Polish Wiktionary usage examples",
45
  "provenance": "Parsed directly from the official Wikimedia dump "
@@ -58,6 +64,8 @@ SOURCES = {
58
  "is_ocr": False,
59
  },
60
  "eurlex": {
 
 
61
  "speakleash_key": "eurlex_corpus",
62
  "pretty": "EUR-Lex (EU legal acts, Polish)",
63
  "license": "CC-BY-4.0",
@@ -71,6 +79,8 @@ SOURCES = {
71
  "is_ocr": False,
72
  },
73
  "parliamentary": {
 
 
74
  "speakleash_key": "PPC_corpus",
75
  "pretty": "Polish Parliamentary Corpus (Sejm/Senat)",
76
  "license": "public-domain (official documents)",
@@ -84,6 +94,9 @@ SOURCES = {
84
  "is_ocr": False,
85
  },
86
  "sejm_api": {
 
 
 
87
  "file_key": "sejm_api",
88
  "pretty": "Sejm API parliamentary speeches (2023 onward)",
89
  "license": "public-domain (official documents)",
@@ -100,6 +113,7 @@ SOURCES = {
100
  "is_ocr": False,
101
  },
102
  "sejm_interpellations": {
 
103
  "file_key": "sejm_interpellations",
104
  "pretty": "Sejm interpellations and written questions (terms 7–10)",
105
  "license": "public-domain (official documents)",
@@ -119,6 +133,8 @@ SOURCES = {
119
  "custom_datasheet": True,
120
  },
121
  "parlamint_pl": {
 
 
122
  "file_key": "parlamint_pl",
123
  "pretty": "ParlaMint-PL (parliamentary debates, 2015-2022)",
124
  "license": "CC-BY-4.0",
@@ -136,6 +152,8 @@ SOURCES = {
136
  "is_ocr": False,
137
  },
138
  "wolne_lektury": {
 
 
139
  "speakleash_key": "wolne_lektury_corpus",
140
  "pretty": "Wolne Lektury (school readings)",
141
  "license": "CC-BY-SA-4.0 / Wolna Sztuka 1.3",
@@ -148,6 +166,8 @@ SOURCES = {
148
  "is_ocr": False,
149
  },
150
  "wikinews": {
 
 
151
  "file_key": "plwikinews", # fetched from Wikimedia dumps, NOT via SpeakLeash
152
  "pretty": "Polish Wikinews",
153
  "license": "CC-BY-2.5",
@@ -163,6 +183,8 @@ SOURCES = {
163
  "is_ocr": False,
164
  },
165
  "wikivoyage": {
 
 
166
  "file_key": "plwikivoyage",
167
  "pretty": "Polish Wikivoyage (travel guides)",
168
  "license": "CC-BY-SA-3.0",
@@ -178,6 +200,8 @@ SOURCES = {
178
  "is_ocr": False,
179
  },
180
  "wikibooks": {
 
 
181
  "file_key": "plwikibooks",
182
  "pretty": "Polish Wikibooks (open textbooks)",
183
  "license": "CC-BY-SA-3.0",
@@ -193,6 +217,8 @@ SOURCES = {
193
  "is_ocr": False,
194
  },
195
  "wikiquote": {
 
 
196
  "file_key": "plwikiquote",
197
  "pretty": "Polish Wikiquote (quotations)",
198
  "license": "CC-BY-SA-3.0",
@@ -208,6 +234,8 @@ SOURCES = {
208
  "is_ocr": False,
209
  },
210
  "eltec_pol": {
 
 
211
  "file_key": "eltec_pol",
212
  "pretty": "ELTeC-pol (European Literary Text Collection, Polish)",
213
  "license": "CC-BY-4.0",
@@ -223,6 +251,8 @@ SOURCES = {
223
  "is_ocr": False,
224
  },
225
  "dziennik_ustaw": {
 
 
226
  "file_key": "dziennik_ustaw", # fetched directly, NOT via SpeakLeash
227
  "pretty": "Dziennik Ustaw + Monitor Polski (Polish primary legislation)",
228
  "license": "public-domain (official documents)",
@@ -240,6 +270,8 @@ SOURCES = {
240
  "is_ocr": False,
241
  },
242
  "govpl": {
 
 
243
  "file_key": "govpl", # fetched directly from gov.pl, NOT via SpeakLeash
244
  "pretty": "gov.pl — Polish government press releases",
245
  "license": "CC-BY-SA-4.0",
@@ -262,6 +294,8 @@ SOURCES = {
262
  "is_ocr": False,
263
  },
264
  "samorzad_gov_pl": {
 
 
265
  "file_key": "samorzad_gov_pl",
266
  "pretty": "samorzad.gov.pl — Polish public-sector institutions",
267
  "license": "CC-BY-SA-4.0",
@@ -285,6 +319,9 @@ SOURCES = {
285
  "custom_datasheet": True,
286
  },
287
  "biblioteka_nauki": {
 
 
 
288
  "file_key": "biblioteka_nauki_pl_corpus",
289
  "pretty": "Biblioteka Nauki",
290
  "license": "per-record upstream license",
@@ -303,6 +340,7 @@ SOURCES = {
303
  "is_ocr": False,
304
  },
305
  "europeana": {
 
306
  # Direct rebuild output from fetch_europeana.py, never the raw
307
  # SpeakLeash Europeana aggregate.
308
  "file_key": "europeana",
@@ -328,6 +366,9 @@ SOURCES = {
328
  "legal review confirms their US status.",
329
  },
330
  "nkjp1m": {
 
 
 
331
  "file_key": "nkjp1m",
332
  "pretty": "The manually annotated 1-million word subcorpus of the National Corpus of Polish",
333
  "license": "CC-BY",
@@ -340,6 +381,8 @@ SOURCES = {
340
  "is_ocr": False,
341
  },
342
  "european_hplt_v3_pl": {
 
 
343
  "file_key": "european_hplt_v3_pl",
344
  "pretty": "HPLT v3.0 Polish (web, top WDS bins 10-5)",
345
  "license": "CC0-1.0",
@@ -371,6 +414,8 @@ SOURCES = {
371
  "custom_datasheet": True,
372
  },
373
  "1000_novels": {
 
 
374
  "speakleash_key": "1000_novels_corpus_CLARIN-PL",
375
  "pretty": "1000 Novels Corpus (CLARIN-PL)",
376
  "license": "CC-BY-4.0",
@@ -388,6 +433,8 @@ SOURCES = {
388
  "custom_datasheet": True,
389
  },
390
  "global_voices": {
 
 
391
  # Data contributed directly (PR #7): no in-repo fetch/build script and a
392
  # hand-authored datasheet. Registered for the source table + release totals;
393
  # the datasheet stays contributor-owned via custom_datasheet.
@@ -407,6 +454,7 @@ SOURCES = {
407
  },
408
  }
409
 
 
410
  ADDED = "2026-07-19"
411
  EXCLUDED = {
412
  "open_subtitles_corpus": "Derivative of copyrighted film/TV dialogue; "
 
16
  # be rebuilt from direct upstream/export scripts rather than blind SpeakLeash pulls.
17
  SOURCES = {
18
  "wikipedia": {
19
+ "added": "2026-07-02",
20
+ "release": "0.2.0",
21
  "speakleash_key": "plwiki",
22
  "pretty": "Polish Wikipedia",
23
  "license": "CC-BY-SA-3.0",
 
30
  "is_ocr": False,
31
  },
32
  "wikisource": {
33
+ "added": "2026-07-02",
34
+ "release": "0.2.0",
35
  "speakleash_key": "plwikisource",
36
  "pretty": "Polish Wikisource",
37
  "license": "CC-BY-SA-3.0",
 
44
  "is_ocr": False,
45
  },
46
  "wiktionary_examples": {
47
+ "added": "2026-07-19",
48
+ "release": None,
49
  "file_key": "wiktionary_examples",
50
  "pretty": "Polish Wiktionary usage examples",
51
  "provenance": "Parsed directly from the official Wikimedia dump "
 
64
  "is_ocr": False,
65
  },
66
  "eurlex": {
67
+ "added": "2026-07-02",
68
+ "release": "0.2.0",
69
  "speakleash_key": "eurlex_corpus",
70
  "pretty": "EUR-Lex (EU legal acts, Polish)",
71
  "license": "CC-BY-4.0",
 
79
  "is_ocr": False,
80
  },
81
  "parliamentary": {
82
+ "added": "2026-07-02",
83
+ "release": "0.2.0",
84
  "speakleash_key": "PPC_corpus",
85
  "pretty": "Polish Parliamentary Corpus (Sejm/Senat)",
86
  "license": "public-domain (official documents)",
 
94
  "is_ocr": False,
95
  },
96
  "sejm_api": {
97
+ "added": "2026-08-22",
98
+ "custom_datasheet": True,
99
+ "release": None,
100
  "file_key": "sejm_api",
101
  "pretty": "Sejm API parliamentary speeches (2023 onward)",
102
  "license": "public-domain (official documents)",
 
113
  "is_ocr": False,
114
  },
115
  "sejm_interpellations": {
116
+ "release": None,
117
  "file_key": "sejm_interpellations",
118
  "pretty": "Sejm interpellations and written questions (terms 7–10)",
119
  "license": "public-domain (official documents)",
 
133
  "custom_datasheet": True,
134
  },
135
  "parlamint_pl": {
136
+ "added": "2026-07-19",
137
+ "release": None,
138
  "file_key": "parlamint_pl",
139
  "pretty": "ParlaMint-PL (parliamentary debates, 2015-2022)",
140
  "license": "CC-BY-4.0",
 
152
  "is_ocr": False,
153
  },
154
  "wolne_lektury": {
155
+ "added": "2026-07-02",
156
+ "release": "0.2.0",
157
  "speakleash_key": "wolne_lektury_corpus",
158
  "pretty": "Wolne Lektury (school readings)",
159
  "license": "CC-BY-SA-4.0 / Wolna Sztuka 1.3",
 
166
  "is_ocr": False,
167
  },
168
  "wikinews": {
169
+ "added": "2026-07-02",
170
+ "release": "0.2.0",
171
  "file_key": "plwikinews", # fetched from Wikimedia dumps, NOT via SpeakLeash
172
  "pretty": "Polish Wikinews",
173
  "license": "CC-BY-2.5",
 
183
  "is_ocr": False,
184
  },
185
  "wikivoyage": {
186
+ "added": "2026-07-02",
187
+ "release": "0.2.0",
188
  "file_key": "plwikivoyage",
189
  "pretty": "Polish Wikivoyage (travel guides)",
190
  "license": "CC-BY-SA-3.0",
 
200
  "is_ocr": False,
201
  },
202
  "wikibooks": {
203
+ "added": "2026-07-02",
204
+ "release": "0.2.0",
205
  "file_key": "plwikibooks",
206
  "pretty": "Polish Wikibooks (open textbooks)",
207
  "license": "CC-BY-SA-3.0",
 
217
  "is_ocr": False,
218
  },
219
  "wikiquote": {
220
+ "added": "2026-07-02",
221
+ "release": "0.2.0",
222
  "file_key": "plwikiquote",
223
  "pretty": "Polish Wikiquote (quotations)",
224
  "license": "CC-BY-SA-3.0",
 
234
  "is_ocr": False,
235
  },
236
  "eltec_pol": {
237
+ "added": "2026-07-02",
238
+ "release": "0.2.0",
239
  "file_key": "eltec_pol",
240
  "pretty": "ELTeC-pol (European Literary Text Collection, Polish)",
241
  "license": "CC-BY-4.0",
 
251
  "is_ocr": False,
252
  },
253
  "dziennik_ustaw": {
254
+ "added": "2026-07-02",
255
+ "release": "0.2.0",
256
  "file_key": "dziennik_ustaw", # fetched directly, NOT via SpeakLeash
257
  "pretty": "Dziennik Ustaw + Monitor Polski (Polish primary legislation)",
258
  "license": "public-domain (official documents)",
 
270
  "is_ocr": False,
271
  },
272
  "govpl": {
273
+ "added": "2026-07-19",
274
+ "release": "0.2.2",
275
  "file_key": "govpl", # fetched directly from gov.pl, NOT via SpeakLeash
276
  "pretty": "gov.pl — Polish government press releases",
277
  "license": "CC-BY-SA-4.0",
 
294
  "is_ocr": False,
295
  },
296
  "samorzad_gov_pl": {
297
+ "added": "2026-08-14",
298
+ "release": "0.2.5",
299
  "file_key": "samorzad_gov_pl",
300
  "pretty": "samorzad.gov.pl — Polish public-sector institutions",
301
  "license": "CC-BY-SA-4.0",
 
319
  "custom_datasheet": True,
320
  },
321
  "biblioteka_nauki": {
322
+ "added": "2026-06-27",
323
+ "custom_datasheet": True,
324
+ "release": "0.2.4",
325
  "file_key": "biblioteka_nauki_pl_corpus",
326
  "pretty": "Biblioteka Nauki",
327
  "license": "per-record upstream license",
 
340
  "is_ocr": False,
341
  },
342
  "europeana": {
343
+ "release": None,
344
  # Direct rebuild output from fetch_europeana.py, never the raw
345
  # SpeakLeash Europeana aggregate.
346
  "file_key": "europeana",
 
366
  "legal review confirms their US status.",
367
  },
368
  "nkjp1m": {
369
+ "added": "2026-06-15",
370
+ "custom_datasheet": True,
371
+ "release": "0.2.3",
372
  "file_key": "nkjp1m",
373
  "pretty": "The manually annotated 1-million word subcorpus of the National Corpus of Polish",
374
  "license": "CC-BY",
 
381
  "is_ocr": False,
382
  },
383
  "european_hplt_v3_pl": {
384
+ "added": "2026-07-14",
385
+ "release": "0.2.3",
386
  "file_key": "european_hplt_v3_pl",
387
  "pretty": "HPLT v3.0 Polish (web, top WDS bins 10-5)",
388
  "license": "CC0-1.0",
 
414
  "custom_datasheet": True,
415
  },
416
  "1000_novels": {
417
+ "added": "2026-07-02",
418
+ "release": "0.2.1",
419
  "speakleash_key": "1000_novels_corpus_CLARIN-PL",
420
  "pretty": "1000 Novels Corpus (CLARIN-PL)",
421
  "license": "CC-BY-4.0",
 
433
  "custom_datasheet": True,
434
  },
435
  "global_voices": {
436
+ "added": "2026-07-15",
437
+ "release": "0.2.3",
438
  # Data contributed directly (PR #7): no in-repo fetch/build script and a
439
  # hand-authored datasheet. Registered for the source table + release totals;
440
  # the datasheet stays contributor-owned via custom_datasheet.
 
454
  },
455
  }
456
 
457
+ # Fallback stamp for a source with no recorded "added" date of its own.
458
  ADDED = "2026-07-19"
459
  EXCLUDED = {
460
  "open_subtitles_corpus": "Derivative of copyrighted film/TV dialogue; "
src/test_verify_paper_claims.py ADDED
@@ -0,0 +1,190 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Self-check for src/verify_paper_claims.py against a synthetic mini-repo."""
3
+
4
+ import json
5
+ import shutil
6
+ import sys
7
+ import tempfile
8
+ from pathlib import Path
9
+
10
+ sys.path.insert(0, str(Path(__file__).resolve().parent))
11
+ import verify_paper_claims as vpc # noqa: E402
12
+
13
+ TEX = r"""
14
+ \begin{figure}
15
+ \addplot[thick] coordinates {(v0.1.0,1.00)(v0.1.1,3.00)};
16
+ \caption{grew from 1.00 billion tokens and 1 sources (v0.1.0).}
17
+ \end{figure}
18
+ \begin{abstract}
19
+ Release v0.1.1 comprises \num{300} documents and \num{3.00}~billion tokens
20
+ from 2 sources.
21
+ \end{abstract}
22
+ \begin{table}
23
+ \begin{tabular}{lrr}
24
+ Version & Tokens (billions) & Sources \\
25
+ v0.1.0 & 1.00 & 1 \\
26
+ v0.1.1 & 3.00 & 2 \\
27
+ \end{tabular}
28
+ \label{tab:s04-wersje}
29
+ \end{table}
30
+ \begin{table}
31
+ \label{tab:s05-zrodla}
32
+ \begin{tabular}{llr}
33
+ Source & License & Tokens [M] \\
34
+ \texttt{alpha} & CC0 & 2000.0 \\
35
+ \texttt{beta\_two} & CC-BY-SA-4.0 & 1000.0 \\
36
+ \textbf{Total} & CC-BY-SA-4.0 & \textbf{$\approx$3000} \\
37
+ \end{tabular}
38
+ \end{table}
39
+ \begin{table}
40
+ \label{tab:s10-licencje}
41
+ \begin{tabular}{llr}
42
+ License family & Representative sources & Tokens [billions] \\
43
+ CC0 & alpha & $2.00$ \\
44
+ CC-BY-SA & beta two & $1.00$ \\
45
+ Total & & $3.00$ \\
46
+ \end{tabular}
47
+ \end{table}
48
+ This work is pinned to commit \texttt{deadbeef}.
49
+ """
50
+
51
+ CHANGELOG = """# Changelog
52
+
53
+ ## v0.1.1 (2026-01-02)
54
+
55
+ - Added `beta_two`.
56
+ - Totals: 2 sources, 300 docs, 3.00B tokens.
57
+ """
58
+
59
+ README = "| `v0.1.0` | previous stable | 100 | 1.00B | Baseline from 1 open/official sources. |\n"
60
+
61
+ GROWTH = """VERSIONS = ["v0.1.0", "v0.1.1"]
62
+ DOCUMENTS = [100, 300]
63
+ TOKENS = [1_000_000_000, 3_000_000_000]
64
+ """
65
+
66
+ STATS = {"alpha": (200, 2_000_000_000), "beta_two": (100, 1_000_000_000)}
67
+
68
+ SOURCES = """SOURCES = {
69
+ "alpha": {"license": "CC0-1.0", "license_spdx": "CC0-1.0", "release": "0.1.0"},
70
+ "beta_two": {"license": "CC-BY-SA-4.0", "license_spdx": "CC-BY-SA-4.0", "release": "0.1.1"},
71
+ }
72
+ """
73
+
74
+
75
+ def registry(root: Path) -> dict:
76
+ return vpc.load_registry(root)
77
+
78
+
79
+ def edit(path: Path, old: str, new: str) -> None:
80
+ text = path.read_text()
81
+ assert old in text, f"fixture no longer contains {old!r}"
82
+ path.write_text(text.replace(old, new))
83
+
84
+
85
+ def build(root: Path) -> None:
86
+ (root / "src").mkdir(parents=True)
87
+ (root / "CHANGELOG.md").write_text(CHANGELOG)
88
+ (root / "README.md").write_text(README)
89
+ (root / "src" / "plot_version_growth.py").write_text(GROWTH)
90
+ (root / "src" / "sources.py").write_text(SOURCES)
91
+ for name, (docs, tokens) in STATS.items():
92
+ directory = root / "data" / name
93
+ directory.mkdir(parents=True)
94
+ (directory / f"{name}.stats.json").write_text(
95
+ json.dumps({"kept": docs, "tokens": tokens, "license": "CC0"})
96
+ )
97
+ (directory / f"{name}.md").write_text(
98
+ f"# {name}\n- **License:** `{registry(root)[name]['license']}`\n"
99
+ )
100
+ (root / "paper.tex").write_text(TEX)
101
+
102
+
103
+ def run(root: Path) -> int:
104
+ return vpc.main(["--tex", str(root / "paper.tex"), "--root", str(root), "--ref", "worktree"])
105
+
106
+
107
+ def scenario(mutate=None) -> int:
108
+ root = Path(tempfile.mkdtemp())
109
+ try:
110
+ build(root)
111
+ if mutate:
112
+ mutate(root)
113
+ return run(root)
114
+ finally:
115
+ shutil.rmtree(root)
116
+
117
+
118
+ def drop_token_digit(root: Path) -> None:
119
+ path = root / "paper.tex"
120
+ path.write_text(path.read_text().replace("2000.0", "2500.0"))
121
+
122
+
123
+ def wrong_source_count(root: Path) -> None:
124
+ path = root / "paper.tex"
125
+ path.write_text(path.read_text().replace("v0.1.1 & 3.00 & 2", "v0.1.1 & 3.00 & ---"))
126
+
127
+
128
+ def add_source(root: Path, release: str | None) -> None:
129
+ directory = root / "data" / "gamma"
130
+ directory.mkdir(parents=True)
131
+ (directory / "gamma.stats.json").write_text(json.dumps({"kept": 1, "tokens": 1, "license": "CC0"}))
132
+ (directory / "gamma.md").write_text("# gamma\n- **License:** `CC0-1.0`\n")
133
+ edit(root / "src" / "sources.py", "}\n",
134
+ f' "gamma": {{"license": "CC0-1.0", "release": {release!r}}},\n}}\n')
135
+
136
+
137
+ def unlisted_source(root: Path) -> None:
138
+ add_source(root, "0.1.1")
139
+
140
+
141
+ def unreleased_source(root: Path) -> None:
142
+ add_source(root, None)
143
+
144
+
145
+ def wrong_license(root: Path) -> None:
146
+ edit(root / "paper.tex", "CC-BY-SA-4.0 & 1000.0", "CC-BY-4.0 & 1000.0")
147
+
148
+
149
+ def datasheet_license_drift(root: Path) -> None:
150
+ edit(root / "data" / "alpha" / "alpha.md", "CC0-1.0", "CC-BY-4.0")
151
+
152
+
153
+ def undefined_footnote(root: Path) -> None:
154
+ edit(root / "paper.tex", "& CC0 &", r"& CC0$^{\dagger}$ &")
155
+
156
+
157
+ def family_value_drift(root: Path) -> None:
158
+ edit(root / "paper.tex", "& beta two & $1.00$", "& beta two & $1.50$")
159
+
160
+
161
+ def incomplete_partition(root: Path) -> None:
162
+ edit(root / "paper.tex", "CC-BY-SA & beta two & $1.00$ \\\\\n", "")
163
+
164
+
165
+ def unfootnoted_dual_license(root: Path) -> None:
166
+ """alpha's content is CC0 but its platform redistributes under CC-BY-4.0."""
167
+ edit(root / "src" / "sources.py",
168
+ '"alpha": {"license": "CC0-1.0", "license_spdx": "CC0-1.0"',
169
+ '"alpha": {"license": "CC0-1.0", "license_spdx": "CC-BY-4.0"')
170
+
171
+
172
+ def stale_growth_series(root: Path) -> None:
173
+ path = root / "src" / "plot_version_growth.py"
174
+ path.write_text(path.read_text().replace("3_000_000_000", "3_100_000_000"))
175
+
176
+
177
+ if __name__ == "__main__":
178
+ assert scenario() == 0, "consistent fixture must pass"
179
+ assert scenario(drop_token_digit) == 1, "per-source token drift must fail"
180
+ assert scenario(wrong_source_count) == 1, "carried-over '---' source count must fail"
181
+ assert scenario(unlisted_source) == 1, "source present in repo but not in Table 3 must fail"
182
+ assert scenario(stale_growth_series) == 1, "growth series vs composition drift must fail"
183
+ assert scenario(unreleased_source) == 0, "a source with release=None is not part of the release"
184
+ assert scenario(wrong_license) == 1, "license in Table 3 vs registry must fail"
185
+ assert scenario(datasheet_license_drift) == 1, "datasheet vs registry license drift must fail"
186
+ assert scenario(undefined_footnote) == 1, "a footnote marker with no definition must fail"
187
+ assert scenario(unfootnoted_dual_license) == 1, "an unfootnoted dual license must fail"
188
+ assert scenario(family_value_drift) == 1, "a wrong license-family subtotal must fail"
189
+ assert scenario(incomplete_partition) == 1, "a license family with no row must fail"
190
+ print("ok")
src/verify_paper_claims.py ADDED
@@ -0,0 +1,637 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ #!/usr/bin/env python3
2
+ """Verify the paper's numeric claims against repository ground truth.
3
+
4
+ Two checks, matching the two things the paper asserts numerically:
5
+
6
+ 1. Release trajectory (Table `tab:s04-wersje` + the hero figure):
7
+ tokens and source counts for v0.2.0 -> v0.2.5.
8
+ Ground truth: `src/plot_version_growth.py` literals (exact docs/tokens)
9
+ and CHANGELOG.md / README.md (source counts per release).
10
+
11
+ 2. Corpus composition (Table `tab:s05-zrodla` + the inline totals):
12
+ per-source token counts, the source set, and the corpus totals.
13
+ Ground truth: `data/*/*.stats.json` at a git ref.
14
+
15
+ Usage:
16
+ python src/verify_paper_claims.py # ref = pin found in the .tex
17
+ python src/verify_paper_claims.py --ref worktree # current working tree
18
+ python src/verify_paper_claims.py --tex paper.tex --ref HEAD
19
+ """
20
+
21
+ from __future__ import annotations
22
+
23
+ import argparse
24
+ import ast
25
+ import json
26
+ import re
27
+ import subprocess
28
+ import sys
29
+ from pathlib import Path
30
+
31
+ ROOT = Path(__file__).resolve().parent.parent
32
+
33
+ # Printed precision in the paper sets the comparison tolerance.
34
+ TOL_BILLIONS = 0.005 # Table 2 / hero figure: 2 decimal places
35
+ TOL_MILLIONS = 0.05 # Table 3: 1 decimal place
36
+
37
+ OK, FAIL, WARN = "OK", "FAIL", "WARN"
38
+
39
+ MARKER = re.compile(r"\$\^\{([^}]*)\}\$")
40
+
41
+
42
+ class Report:
43
+ def __init__(self) -> None:
44
+ self.rows: list[tuple[str, str, str]] = []
45
+
46
+ def add(self, status: str, label: str, detail: str = "") -> None:
47
+ self.rows.append((status, label, detail))
48
+
49
+ def num(self, label: str, claimed, actual, tol: float = 0.0) -> None:
50
+ if claimed is None:
51
+ self.add(WARN, label, "not found in .tex")
52
+ elif abs(claimed - actual) <= tol:
53
+ self.add(OK, label, f"{claimed} == {actual}")
54
+ else:
55
+ self.add(FAIL, label, f"paper {claimed} != repo {actual}")
56
+
57
+ def section(self, title: str) -> None:
58
+ self.rows.append(("", title, ""))
59
+
60
+ def render(self) -> int:
61
+ width = max(len(label) for _, label, _ in self.rows) + 2
62
+ for status, label, detail in self.rows:
63
+ if not status:
64
+ print(f"\n=== {label} ===")
65
+ else:
66
+ print(f" [{status:<4}] {label:<{width}} {detail}")
67
+ failures = sum(1 for s, _, _ in self.rows if s == FAIL)
68
+ warnings = sum(1 for s, _, _ in self.rows if s == WARN)
69
+ print(f"\n{failures} failure(s), {warnings} warning(s)")
70
+ return 1 if failures else 0
71
+
72
+
73
+ # --------------------------------------------------------------------------
74
+ # .tex parsing
75
+ # --------------------------------------------------------------------------
76
+
77
+ def _unescape(name: str) -> str:
78
+ return name.replace("\\_", "_").replace("\\&", "&")
79
+
80
+
81
+ def _table_body(tex: str, label: str) -> str:
82
+ """Return the body of the table environment carrying \\label{label}."""
83
+ for block in re.findall(r"\\begin\{table\}.*?\\end\{table\}", tex, re.S):
84
+ if f"\\label{{{label}}}" in block:
85
+ return block
86
+ raise LookupError(f"no table with \\label{{{label}}} in the .tex")
87
+
88
+
89
+ def parse_hero(tex: str) -> dict[str, float]:
90
+ """Version -> tokens (billions) from the pgfplots hero figure coordinates."""
91
+ coords = re.search(r"\\addplot\[[^\]]*\]\s*coordinates\s*\{(.*?)\}", tex, re.S)
92
+ if not coords:
93
+ return {}
94
+ return {v: float(t) for v, t in re.findall(r"\((v[\d.]+),\s*([\d.]+)\)", coords.group(1))}
95
+
96
+
97
+ def parse_trajectory_table(tex: str) -> dict[str, tuple[float, str]]:
98
+ """Version -> (tokens in billions, source-count cell verbatim)."""
99
+ body = _table_body(tex, "tab:s04-wersje")
100
+ rows = re.findall(r"^\s*(v[\d.]+)\s*&\s*([\d.]+)\s*&\s*(.+?)\s*\\\\", body, re.M)
101
+ return {v: (float(tok), src.strip()) for v, tok, src in rows}
102
+
103
+
104
+ def parse_source_table(tex: str) -> tuple[dict[str, float], float | None]:
105
+ """(source name -> tokens in millions, declared total in millions)."""
106
+ body = _table_body(tex, "tab:s05-zrodla")
107
+ sources = {
108
+ _unescape(name): float(tok)
109
+ for name, tok in re.findall(
110
+ r"\\texttt\{([^}]+)\}\s*&.*?&\s*([\d.]+)\s*\\\\", body
111
+ )
112
+ }
113
+ total_row = re.search(r"\\textbf\{Total\}.*?&.*?&\s*(.+?)\s*\\\\", body, re.S)
114
+ total = None
115
+ if total_row:
116
+ digits = re.search(r"([\d.]+)", total_row.group(1))
117
+ if digits:
118
+ total = float(digits.group(1))
119
+ return sources, total
120
+
121
+
122
+ def parse_license_rows(tex: str) -> dict[str, tuple[str, set[str]]]:
123
+ """Table 3 source rows -> (license cell without its markers, marker names)."""
124
+ body = _table_body(tex, "tab:s05-zrodla").split(r"\begin{tabular}")[-1]
125
+ rows = {}
126
+ for name, cell in re.findall(r"\\texttt\{([^}]+)\}\s*&(.*?)&\s*[\d.]+\s*\\\\", body):
127
+ rows[_unescape(name)] = (MARKER.sub("", cell).strip(), _markers(cell))
128
+ return rows
129
+
130
+
131
+ def parse_footnote_defs(tex: str) -> set[str]:
132
+ """Markers the Table 3 caption actually explains (a definition is followed by ~)."""
133
+ caption = _table_body(tex, "tab:s05-zrodla").split(r"\begin{tabular}")[0]
134
+ return {m for cell in re.findall(r"(\$\^\{[^}]*\}\$)~", caption) for m in _markers(cell)}
135
+
136
+
137
+ def parse_total_license(tex: str) -> str | None:
138
+ body = _table_body(tex, "tab:s05-zrodla")
139
+ row = re.search(r"\\textbf\{Total\}\s*&\s*(.*?)\s*&", body)
140
+ return MARKER.sub("", row.group(1)).strip() if row else None
141
+
142
+
143
+ def _markers(cell: str) -> set[str]:
144
+ return {m for group in MARKER.findall(cell) for m in re.findall(r"\\([a-zA-Z]+)", group)}
145
+
146
+
147
+ def parse_family_table(tex: str) -> tuple[dict[str, float], float | None]:
148
+ """Table 5 -> (license family -> tokens in billions, declared total)."""
149
+ body = _table_body(tex, "tab:s10-licencje").split(r"\begin{tabular}")[-1]
150
+ families, total = {}, None
151
+ for label, value in re.findall(r"^\s*([^&\\]+?)\s*&[^&]*&\s*\$?([\d.]+)\$?\s*\\\\",
152
+ body, re.M):
153
+ if label.strip().lower() == "total":
154
+ total = float(value)
155
+ elif license_family(label):
156
+ families[license_family(label)] = float(value)
157
+ return families, total
158
+
159
+
160
+ def license_family(text: str) -> str | None:
161
+ """The Table 5 partition: like norm_license, but CC0 and public domain split.
162
+
163
+ norm_license folds them together because the source table prints either for
164
+ the same source; Table 5 reports them as separate families, so a source's
165
+ own wording decides which side it lands on.
166
+ """
167
+ t = re.sub(r"[`_]", " ", text).upper()
168
+ if re.search(r"PER-RECORD|PER-DOCUMENT|MIXED", t):
169
+ return "Per-record (mixed)"
170
+ if re.search(r"PUBLIC[- ]DOMAIN", t):
171
+ return "Public domain"
172
+ if "CC0" in t:
173
+ return "CC0"
174
+ if re.search(r"CC[- ]BY[- ]SA", t):
175
+ return "CC-BY-SA"
176
+ if re.search(r"CC[- ]BY", t):
177
+ return "CC-BY"
178
+ return None
179
+
180
+
181
+ def parse_inline_claims(tex: str, current: str | None = None) -> dict:
182
+ """Claims the prose/captions make about the *current release* totals.
183
+
184
+ A claim is skipped when its surrounding sentence scopes it to a different
185
+ release or to a partition subtotal, so only whole-release assertions are
186
+ compared against the current release's ground truth.
187
+ """
188
+ CONTEXT = 160
189
+ SUBTOTAL = re.compile(r"share-alike|partitions?\b")
190
+ VERSION = re.compile(r"v\d+\.\d+\.\d+")
191
+
192
+ def sweep(pattern: str, cast):
193
+ out = set()
194
+ for match in re.finditer(pattern, tex):
195
+ window = tex[max(0, match.start() - CONTEXT):match.end() + CONTEXT]
196
+ if SUBTOTAL.search(window):
197
+ continue
198
+ if any(v != current for v in VERSION.findall(window)):
199
+ continue
200
+ raw = next(g for g in match.groups() if g)
201
+ out.add(cast(raw))
202
+ return out
203
+
204
+ docs = sweep(
205
+ r"(?:\\num\{(\d+)\}|(\d[\d\\,]*\d))\s*~?\s*documents",
206
+ lambda raw: int(re.sub(r"[^\d]", "", raw)),
207
+ )
208
+ # ponytail: floor keeps small unrelated "N documents" counts (held-out set
209
+ # sizes) out; scope by section if the paper grows other six-figure claims.
210
+ docs = {value for value in docs if value >= 100_000}
211
+
212
+ tokens = sweep(
213
+ r"(?:\\num\{([\d.]+)\}|\$([\d.]+)\$|([\d.]+))\s*~?\s*billion\s+tokens",
214
+ float,
215
+ )
216
+ sources = sweep(r"\b(\d+)\s+sources\b", int)
217
+
218
+ pin = re.search(r"commit\s+\\texttt\{([0-9a-f]{7,40})\}", tex)
219
+ return {
220
+ "documents": docs,
221
+ "tokens_billions": tokens,
222
+ "sources": sources,
223
+ "pin": pin.group(1) if pin else None,
224
+ }
225
+
226
+
227
+ # --------------------------------------------------------------------------
228
+ # Ground truth
229
+ # --------------------------------------------------------------------------
230
+
231
+ def load_growth_series(root: Path) -> dict[str, tuple[int, int]]:
232
+ """Version -> (documents, tokens) from src/plot_version_growth.py literals."""
233
+ path = root / "src" / "plot_version_growth.py"
234
+ tree = ast.parse(path.read_text(encoding="utf-8"))
235
+ found: dict[str, list] = {}
236
+ for node in tree.body:
237
+ if isinstance(node, ast.Assign) and len(node.targets) == 1:
238
+ target = node.targets[0]
239
+ if isinstance(target, ast.Name) and target.id in {"VERSIONS", "DOCUMENTS", "TOKENS"}:
240
+ found[target.id] = ast.literal_eval(node.value)
241
+ missing = {"VERSIONS", "DOCUMENTS", "TOKENS"} - found.keys()
242
+ if missing:
243
+ raise LookupError(f"{path} is missing {sorted(missing)}")
244
+ return dict(zip(found["VERSIONS"], zip(found["DOCUMENTS"], found["TOKENS"])))
245
+
246
+
247
+ def load_registry(root: Path) -> dict[str, dict]:
248
+ """The SOURCES table from src/sources.py, read without importing the module.
249
+
250
+ Always read from the working tree: the registry defines what a release *is*,
251
+ independently of the ref the parquet stats are read from.
252
+ """
253
+ path = root / "src" / "sources.py"
254
+ if not path.exists():
255
+ return {}
256
+ for node in ast.parse(path.read_text(encoding="utf-8")).body:
257
+ if (isinstance(node, ast.Assign) and len(node.targets) == 1
258
+ and getattr(node.targets[0], "id", None) == "SOURCES"):
259
+ return ast.literal_eval(node.value)
260
+ return {}
261
+
262
+
263
+ def load_release_map(root: Path) -> dict[str, str | None]:
264
+ """Source name -> the release it entered.
265
+
266
+ Empty when the registry is absent or predates the `release` field, in which
267
+ case the CHANGELOG prose stays the only ground truth.
268
+ """
269
+ sources = load_registry(root)
270
+ if not any("release" in cfg for cfg in sources.values()):
271
+ return {}
272
+ return {name: cfg.get("release") for name, cfg in sources.items()}
273
+
274
+
275
+ def released_at(releases: dict[str, str | None], version: str) -> set[str]:
276
+ """Sources belonging to `version`, i.e. admitted at or before it."""
277
+ ceiling = _version_key(version)
278
+ return {
279
+ name for name, release in releases.items()
280
+ if release and _version_key(release) <= ceiling
281
+ }
282
+
283
+
284
+ def _version_key(version: str) -> tuple[int, ...]:
285
+ return tuple(int(part) for part in version.lstrip("v").split("."))
286
+
287
+
288
+ def _changelog_sections(root: Path) -> dict[str, str]:
289
+ text = (root / "CHANGELOG.md").read_text(encoding="utf-8")
290
+ parts = re.split(r"^## (v[\d.]+[\w.-]*)", text, flags=re.M)
291
+ return dict(zip(parts[1::2], parts[2::2]))
292
+
293
+
294
+ def load_declared_sources(
295
+ root: Path, base_version: str | None = None
296
+ ) -> tuple[dict[str, int | None], dict[str, list[str]]]:
297
+ """Version -> declared source count (None when unrecorded), and sources added."""
298
+ sections = _changelog_sections(root)
299
+ counts: dict[str, int | None] = {}
300
+ added: dict[str, list[str]] = {}
301
+ for version, body in sections.items():
302
+ match = re.search(r"\b(\d+)[-\s]sources?\b", body)
303
+ counts[version] = int(match.group(1)) if match else None
304
+ added[version] = re.findall(
305
+ r"Added (?:the )?(?:community[- ]contributed |community contribution |reviewed )?`([\w.]+)`",
306
+ body,
307
+ )
308
+ # The first release predates the CHANGELOG; the README release table records it.
309
+ if base_version and counts.get(base_version) is None:
310
+ readme = (root / "README.md").read_text(encoding="utf-8")
311
+ base = re.search(
312
+ rf"`{re.escape(base_version)}`.*?(\d+) open/official sources", readme, re.S
313
+ )
314
+ if base:
315
+ counts[base_version] = int(base.group(1))
316
+ added.setdefault(base_version, [])
317
+ return counts, added
318
+
319
+
320
+ def load_stats(root: Path, ref: str) -> dict[str, dict]:
321
+ """Source name -> parquet stats, read from a git ref or the working tree."""
322
+ stats: dict[str, dict] = {}
323
+ if ref == "worktree":
324
+ paths = sorted(root.glob("data/*/*.stats.json"))
325
+ for path in paths:
326
+ stats[path.name.removesuffix(".stats.json")] = json.loads(
327
+ path.read_text(encoding="utf-8")
328
+ )
329
+ return stats
330
+
331
+ listing = subprocess.run(
332
+ ["git", "-C", str(root), "ls-tree", "-r", "--name-only", ref, "--", "data/"],
333
+ capture_output=True, text=True, check=True,
334
+ ).stdout.splitlines()
335
+ for rel in sorted(p for p in listing if p.endswith(".stats.json")):
336
+ blob = subprocess.run(
337
+ ["git", "-C", str(root), "show", f"{ref}:{rel}"],
338
+ capture_output=True, text=True, check=True,
339
+ ).stdout
340
+ stats[Path(rel).name.removesuffix(".stats.json")] = json.loads(blob)
341
+ return stats
342
+
343
+
344
+ # --------------------------------------------------------------------------
345
+ # Checks
346
+ # --------------------------------------------------------------------------
347
+
348
+ def check_trajectory(report: Report, tex: str, root: Path) -> None:
349
+ report.section("Check 1 - release trajectory v0.2.0 -> v0.2.5 (Table 2 + hero figure)")
350
+
351
+ table = parse_trajectory_table(tex)
352
+ hero = parse_hero(tex)
353
+ series = load_growth_series(root)
354
+ declared, added = load_declared_sources(root, base_version=sorted(series)[0])
355
+ releases = load_release_map(root)
356
+ if releases:
357
+ for version in sorted(series):
358
+ counted = len(released_at(releases, version))
359
+ if declared.get(version) is None:
360
+ declared[version] = counted
361
+ elif declared[version] != counted:
362
+ report.add(FAIL, f"{version} registry vs CHANGELOG",
363
+ f"src/sources.py counts {counted}, CHANGELOG says {declared[version]}")
364
+
365
+ missing = sorted(set(series) - set(table))
366
+ extra = sorted(set(table) - set(series))
367
+ if missing or extra:
368
+ report.add(FAIL, "table covers every release",
369
+ f"missing {missing}, unexpected {extra}")
370
+ else:
371
+ report.add(OK, "table covers every release", f"{len(table)} rows")
372
+
373
+ for version in sorted(series):
374
+ docs, tokens = series[version]
375
+ actual_b = tokens / 1e9
376
+ claimed_b = table.get(version, (None, None))[0]
377
+ report.num(f"{version} tokens (B, Table 2)", claimed_b, round(actual_b, 6), TOL_BILLIONS)
378
+ report.num(f"{version} tokens (B, hero figure)", hero.get(version), round(actual_b, 6), TOL_BILLIONS)
379
+
380
+ # Source counts. "---" in the table means "unchanged from the previous release".
381
+ previous: int | None = None
382
+ for version in sorted(series):
383
+ cell = table.get(version, (None, "?"))[1]
384
+ if cell == "---":
385
+ claimed = previous
386
+ shown = f"--- (carries {previous})"
387
+ else:
388
+ digits = re.search(r"\d+", cell or "")
389
+ claimed = int(digits.group()) if digits else None
390
+ shown = str(claimed)
391
+ previous = claimed
392
+
393
+ truth = declared.get(version)
394
+ label = f"{version} source count"
395
+ if truth is None:
396
+ report.add(WARN, label, f"paper {shown}; no count recorded in CHANGELOG/README")
397
+ elif claimed == truth:
398
+ report.add(OK, label, f"{shown} == {truth}")
399
+ else:
400
+ report.add(FAIL, label, f"paper {shown} != repo {truth}")
401
+
402
+ if releases:
403
+ return # the registry already resolves every release exactly
404
+
405
+ # Back-fill unrecorded releases from the next release's "Added `source`" bullets.
406
+ versions = sorted(series)
407
+ for index, version in enumerate(versions[:-1]):
408
+ if declared.get(version) is not None:
409
+ continue
410
+ following = versions[index + 1]
411
+ if declared.get(following) is None:
412
+ continue
413
+ inferred = declared[following] - len(added.get(following, []))
414
+ report.add(
415
+ WARN,
416
+ f"{version} source count (derived)",
417
+ f"{following} declares {declared[following]} and adds "
418
+ f"{len(added.get(following, []))} -> {version} = {inferred}",
419
+ )
420
+
421
+
422
+ def check_composition(report: Report, tex: str, root: Path, ref: str) -> None:
423
+ report.section(f"Check 2 - corpus composition and totals (Table 3, ref={ref})")
424
+
425
+ table, table_total = parse_source_table(tex)
426
+ current = sorted(load_growth_series(root))[-1]
427
+ claims = parse_inline_claims(tex, current=current)
428
+ stats = load_stats(root, ref)
429
+ releases = load_release_map(root)
430
+ in_release = released_at(releases, current) if releases else set(stats)
431
+
432
+ absent = sorted(set(table) - set(stats))
433
+ unlisted = sorted((set(stats) & in_release) - set(table))
434
+ unreleased = sorted(set(stats) - in_release)
435
+ listed_but_unreleased = sorted(set(table) & set(unreleased))
436
+ if absent:
437
+ report.add(FAIL, "every listed source exists at ref", f"missing from repo: {absent}")
438
+ else:
439
+ report.add(OK, "every listed source exists at ref", f"{len(table)} sources")
440
+ if unlisted:
441
+ report.add(FAIL, f"every {current} source is in Table 3",
442
+ f"released but missing from Table 3: {unlisted}")
443
+ else:
444
+ report.add(OK, f"every {current} source is in Table 3", "")
445
+ if listed_but_unreleased:
446
+ report.add(FAIL, "Table 3 lists only released sources",
447
+ f"not in any release: {listed_but_unreleased}")
448
+ elif unreleased:
449
+ report.add(OK, "Table 3 lists only released sources",
450
+ f"correctly excludes {unreleased}")
451
+
452
+ listed = [name for name in table if name in stats]
453
+ for name in sorted(listed, key=lambda n: -stats[n]["tokens"]):
454
+ report.num(f"{name} tokens (M)", table[name], round(stats[name]["tokens"] / 1e6, 6), TOL_MILLIONS)
455
+
456
+ sum_tokens = sum(stats[name]["tokens"] for name in listed)
457
+ sum_docs = sum(stats[name]["kept"] for name in listed)
458
+
459
+ report.num("Table 3 rows sum to its Total row (M)",
460
+ round(sum(table[n] for n in listed), 6),
461
+ round(sum_tokens / 1e6, 6), TOL_MILLIONS * len(listed))
462
+ report.num("Table 3 Total row (M)", table_total, round(sum_tokens / 1e6, 6), 1.0)
463
+
464
+ for value in sorted(claims["documents"]):
465
+ report.num(f"prose claim: {value} documents", value, sum_docs)
466
+ for value in sorted(claims["tokens_billions"]):
467
+ report.num(f"prose claim: {value}B tokens", value, round(sum_tokens / 1e9, 6), TOL_BILLIONS)
468
+ for value in sorted(claims["sources"]):
469
+ report.num(f"prose claim: {value} sources", value, len(listed))
470
+
471
+ if unreleased:
472
+ report.add(
473
+ WARN, "totals over ALL sources at ref",
474
+ f"{sum(s['kept'] for s in stats.values())} docs / "
475
+ f"{sum(s['tokens'] for s in stats.values()) / 1e9:.3f}B tokens "
476
+ f"across {len(stats)} sources",
477
+ )
478
+
479
+ # Anchor: the trajectory series' last point must equal the composition sum.
480
+ series = load_growth_series(root)
481
+ last = sorted(series)[-1]
482
+ report.num(f"{last} tokens in growth series == Table 3 sum",
483
+ series[last][1], sum_tokens)
484
+ report.num(f"{last} documents in growth series == Table 3 sum",
485
+ series[last][0], sum_docs)
486
+
487
+
488
+ def norm_license(text: str) -> tuple[str, str | None]:
489
+ """Canonical (family, version) so paper, registry and datasheet can be compared.
490
+
491
+ Public domain and CC0 are one family: SPDX spells "official documents are
492
+ outside copyright" as CC0-1.0, and the paper prints both as PD. A version of
493
+ None means the text states no version, and compares equal to any version --
494
+ the paper abbreviates nkjp1m's CC-BY-4.0 to "CC-BY" on purpose.
495
+ """
496
+ t = re.sub(r"[`_]", " ", text).upper()
497
+ if re.search(r"PER-RECORD|PER-DOCUMENT|MIXED", t):
498
+ return ("MIXED", None)
499
+ if re.search(r"CC0|PUBLIC[- ]DOMAIN|\bPD\b", t):
500
+ return ("PD/CC0", None)
501
+ match = re.search(r"CC[- ]BY(?:[- ](SA))?(?:[- ](\d+\.\d+))?", t)
502
+ if match:
503
+ return ("CC-BY-SA" if match.group(1) else "CC-BY", match.group(2))
504
+ return (t.strip(), None)
505
+
506
+
507
+ def compatible(a: tuple[str, str | None], b: tuple[str, str | None]) -> bool:
508
+ """Same family, and same version unless one side states none."""
509
+ return a[0] == b[0] and (a[1] is None or b[1] is None or a[1] == b[1])
510
+
511
+
512
+ def load_datasheet_license(root: Path, name: str) -> str | None:
513
+ """The datasheet's declared License line (working tree, not the pinned ref)."""
514
+ path = root / "data" / name / f"{name}.md"
515
+ if not path.exists():
516
+ return None
517
+ match = re.search(r"\*\*License:\*\*\s*(.+)", path.read_text(encoding="utf-8"))
518
+ return match.group(1).strip() if match else None
519
+
520
+
521
+ def check_families(report: Report, tex: str, root: Path, ref: str) -> None:
522
+ """Table 5 must be a complete partition of Table 3 by license family."""
523
+ paper, paper_total = parse_family_table(tex)
524
+ registry, stats = load_registry(root), load_stats(root, ref)
525
+ releases = load_release_map(root)
526
+ if not (paper and registry and releases):
527
+ report.add(WARN, "license family table", "no family table or no registry to partition")
528
+ return
529
+
530
+ repo: dict[str, int] = {}
531
+ unclassified = []
532
+ for name in released_at(releases, sorted(load_growth_series(root))[-1]):
533
+ cfg = registry[name]
534
+ # The row prints the content's own status; SPDX carries the CC0 wording.
535
+ family = license_family(str(cfg.get("license", ""))) or license_family(str(cfg.get("license_spdx") or ""))
536
+ if family is None:
537
+ unclassified.append(name)
538
+ else:
539
+ repo[family] = repo.get(family, 0) + stats[name]["tokens"]
540
+
541
+ if unclassified:
542
+ report.add(FAIL, "every source has a license family", f"unclassifiable: {sorted(unclassified)}")
543
+ if set(paper) != set(repo):
544
+ report.add(FAIL, "Table 5 families match the corpus",
545
+ f"paper {sorted(paper)} != repo {sorted(repo)}")
546
+ else:
547
+ report.add(OK, "Table 5 families match the corpus", f"{len(repo)} families, no source left over")
548
+
549
+ for family in sorted(repo, key=lambda f: -repo[f]):
550
+ if family in paper:
551
+ report.num(f"{family} tokens (B)", paper[family], round(repo[family] / 1e9, 6), TOL_BILLIONS)
552
+ report.num("Table 5 Total row (B)", paper_total, round(sum(repo.values()) / 1e9, 6), TOL_BILLIONS)
553
+
554
+
555
+ def check_licenses(report: Report, tex: str, root: Path) -> None:
556
+ report.section("Check 3 - licenses, footnotes and license families")
557
+
558
+ rows = parse_license_rows(tex)
559
+ registry = load_registry(root)
560
+ if not registry:
561
+ report.add(WARN, "license audit", "src/sources.py carries no registry to compare against")
562
+ return
563
+
564
+ for name in sorted(rows, key=lambda n: n.lower()):
565
+ paper, markers = rows[name]
566
+ cfg = registry.get(name)
567
+ if cfg is None:
568
+ report.add(FAIL, f"{name} in registry", "listed in Table 3 but absent from src/sources.py")
569
+ continue
570
+
571
+ declared = norm_license(str(cfg.get("license", "")))
572
+ spdx = norm_license(str(cfg.get("license_spdx") or cfg.get("license", "")))
573
+ printed = norm_license(paper)
574
+
575
+ if compatible(printed, declared) or compatible(printed, spdx):
576
+ report.add(OK, f"{name} license", f"{paper} == {cfg.get('license')}")
577
+ else:
578
+ report.add(FAIL, f"{name} license",
579
+ f"paper {paper!r} != registry {cfg.get('license')!r} / {cfg.get('license_spdx')!r}")
580
+
581
+ sheet = load_datasheet_license(root, name)
582
+ if sheet is None:
583
+ report.add(WARN, f"{name} datasheet license", "no **License:** line in the datasheet")
584
+ elif not compatible(norm_license(sheet), declared):
585
+ report.add(FAIL, f"{name} datasheet license",
586
+ f"datasheet {sheet!r} != registry {cfg.get('license')!r}")
587
+
588
+ # A source whose content license differs from its redistribution license
589
+ # states only half the story in a one-line cell; the footnote carries the rest.
590
+ if not compatible(declared, spdx) and not markers:
591
+ report.add(FAIL, f"{name} dual license is footnoted",
592
+ f"content {cfg.get('license')!r} but redistributed as "
593
+ f"{cfg.get('license_spdx')!r}; the row prints only {paper!r}")
594
+
595
+ used = {m for _, markers in rows.values() for m in markers}
596
+ defined = parse_footnote_defs(tex)
597
+ if used - defined:
598
+ report.add(FAIL, "every footnote marker is defined", f"used but not explained: {sorted(used - defined)}")
599
+ elif defined - used:
600
+ report.add(FAIL, "every footnote is used", f"explained but on no row: {sorted(defined - used)}")
601
+ else:
602
+ report.add(OK, "footnote markers", f"{len(used)} defined and used: {sorted(used)}")
603
+
604
+ total = parse_total_license(tex)
605
+ strongest = [n for n, (lic, _) in rows.items() if norm_license(lic)[0] == "CC-BY-SA"]
606
+ if total is None:
607
+ report.add(WARN, "Total row license", "no license cell in the Total row")
608
+ elif strongest and norm_license(total)[0] != "CC-BY-SA":
609
+ report.add(FAIL, "Total row license", f"{total} does not inherit the share-alike of {sorted(strongest)}")
610
+ else:
611
+ report.add(OK, "Total row license", f"{total} inherits share-alike from {len(strongest)} sources")
612
+
613
+
614
+ def main(argv: list[str] | None = None) -> int:
615
+ parser = argparse.ArgumentParser(description=__doc__,
616
+ formatter_class=argparse.RawDescriptionHelpFormatter)
617
+ parser.add_argument("--tex", default=str(ROOT / "file.tex"), help="paper source")
618
+ parser.add_argument("--root", default=str(ROOT), help="repository root")
619
+ parser.add_argument("--ref", default=None,
620
+ help="git ref for data/*.stats.json, or 'worktree' "
621
+ "(default: the commit the .tex pins)")
622
+ args = parser.parse_args(argv)
623
+
624
+ root = Path(args.root).resolve()
625
+ tex = Path(args.tex).read_text(encoding="utf-8")
626
+ ref = args.ref or parse_inline_claims(tex)["pin"] or "HEAD"
627
+
628
+ report = Report()
629
+ check_trajectory(report, tex, root)
630
+ check_composition(report, tex, root, ref)
631
+ check_licenses(report, tex, root)
632
+ check_families(report, tex, root, ref)
633
+ return report.render()
634
+
635
+
636
+ if __name__ == "__main__":
637
+ sys.exit(main())