| Arm | Checkpoint | Prompt |
|---|---|---|
| B1 | end of stage 1 | raw |
| B2 | end of midtraining | raw |
| B3 | final base | raw |
| N1 | Gen-QA mix from 2T | raw |
| N2 | math-code-thinking mix from 2T | raw |
| N3 | Round 5 mix from 2T | raw |
| P1 | SFT | chat template |
| P2 | DPO | chat template |
| P3 | final instruct | chat template |
Study 4: where Olmo 3’s contrastive frames enter its training
englishRegisterStudy — results from the pre-registered analysis
1 What this measures
Studies 2 and 3 of this project found, in analyses registered as exploratory, that Olmo 3 7B writes the contrastive frames, it’s not X, it’s Y and its relatives, at several times the rate of the web text it was pretrained on, and that the instruction-tuned model writes them more often again. Study 4 asks where in the model’s training that excess enters. It was registered at https://osf.io/d79u4 on 29 September 2026, before any of its text was generated or any frame counted in its data.
Olmo 3 was trained in six stages: stage-1 pretraining, midtraining, long-context training, supervised fine-tuning (SFT), preference tuning (DPO) and reinforcement learning with verifiable rewards (RLVR). Ai2 published a checkpoint after each stage and the data each stage trained on (Team Olmo, 2025). Study 4 counts the frames in the output of checkpoints from each stage and in the text four of the stages trained on.
Each checkpoint wrote one post for each of the 4,000 titles Study 2 used, drawn from GloWbE’s blog pages (Davies, 2013) and prompted with Write a blog post titled: and the title, at temperature 1.0 and top-p 1.0 with at most 1,024 new tokens. The base-model checkpoints received the prompt as raw text and the three post-trained ones through their chat template, which adds a default system prompt, as in Study 2. N1 to N3 are three midtraining runs of the same length from one stage-1 checkpoint at 2 trillion tokens, which Ai2 published as branches of the base model, each with a different data mix. In the Olmo 3 report’s description, the Gen-QA mix “increases proportions of web, QA, and instruction data while omitting math, code, and thinking”, the math-code-thinking mix “increases proportions of math, code, and thinking data while omitting QA and instruction data” and keeps web, and Round 5 is the final mix (Team Olmo, 2025).
Four samples of training text are measured:
- Stage 1: Study 3’s registered sample of
common_crawl, the web portion of the stage-1 mix, 200 files and 200.4 million words. - Midtraining: 240.0 million words from 1,015 files of the Dolmino midtraining mix, allocated across its 24 sources by their shares of its tokens, so the pooled sample stands for the mix.
- SFT: the assistant turns of all 2,152,112 conversations in the SFT set, 411.3 million words.
- DPO: the final response on each side of all 259,922 preference pairs, 84.1 million words chosen and 77.9 million rejected.
The long-context stage’s data, two-thirds of it midtraining data, is not counted. The RLVR data holds prompts and reference answers but none of the model’s responses, so RLVR is measured on the output side only.
1.1 Two outcomes
The family is the four contrastive frames F06 to F09, it 's not * , it 's, it 's not just, not just * but and is n't just, as a rate per word. Each source’s rate is its total hits over its total words, with an interval from the variation between its units, whether posts, files, conversations or pairs (R/ratio.R). The composite is all fifteen frozen frames through the model Studies 2 and 3 used, in which a source’s effect is its frame rates averaged over frames on the log scale, so that it means what it meant in Studies 1 to 3.
Studies 2 and 3 summarised the contrastive family over the frames they shared with Study 1, F07 to F10. F10, not * , but *, is a partial frame, and Study 4 registered the family without it, adding it only as a sensitivity check. The two sets give different multiples for the same text. The final base model’s rate is 5.52 times the stage-1 sample’s over F07 to F10, the figure behind the registration’s “about five times”, and 8.46 times over F06 to F09. F10 runs near the stage-1 sample’s rate in every base-model checkpoint, so including it pulls the multiple towards 1.
1.2 Nine hypotheses
- H1. The end-of-stage-1 checkpoint (B1) uses the frames more than its stage-1 input.
- H2. The midtraining text uses them more than the stage-1 text.
- H3. The end-of-midtraining checkpoint (B2) uses them more than B1.
- H4. The Gen-QA checkpoint (N1) and the math-code-thinking checkpoint (N2) differ, in either direction.
- H5. The SFT responses use them more than the final base model (B3).
- H6. The SFT checkpoint (P1) uses them more than B3.
- H7. Chosen DPO responses use them more than rejected responses to the same prompts.
- H8. The DPO checkpoint (P2) uses them more than P1.
- H9. The final instruct model (P3) uses them more than P2.
Each is tested on both outcomes, eighteen tests with Holm’s correction. A hypothesis is supported on an outcome when its ratio lies in the predicted direction with an adjusted p below 0.05, and a ratio in the other direction with an adjusted p below 0.05 is evidence against it.
Earlier work measured post-training’s effect on style from output alone. Reinhart et al. (2025) compared Llama 3 and GPT-4o with human writing on Biber’s grammatical and rhetorical features and found the differences larger for instruction-tuned models than for base models, and Li et al. (2026) compared OLMo 32B’s base, SFT, DPO and RLVR checkpoints with human fiction and found that post-training compresses stylistic variation. Study 4 adds the stages before post-training and the training text of each stage.
1.3 Deviations
Four are recorded in docs/deviations.md, and none changes the design:
- S4-D1: the protocol’s directory pattern for one midtraining source matched nothing and is read as the one pattern that fits.
- S4-D2: the registration gave the DPO set as 260,000 pairs, the dataset card’s rounded description. It holds 259,922, and every pair was counted.
- S4-D3: SFT conversations are identified by file and row, where the protocol names each conversation’s
id. - S4-D4: 353 of the 2,446 documents drawn from the midtraining mix’s olmOCR source were redacted after training and are counted as they stand, as in Study 3.
S4-D1, S4-D2 and S4-D4 were filed as an update to the registration, approved before the analysis ran.
2 The generated text
| Arm | Words | Words a post | Reached 1,024 tokens | No words |
|---|---|---|---|---|
| B1 | 1,873,609 | 468 | 1,596 (40%) | 53 |
| B2 | 1,836,634 | 459 | 1,384 (35%) | 84 |
| B3 | 2,408,514 | 602 | 1,479 (37%) | 3 |
| N1 | 1,891,439 | 473 | 1,396 (35%) | 81 |
| N2 | 2,014,023 | 504 | 1,652 (41%) | 3 |
| N3 | 1,862,425 | 466 | 1,414 (35%) | 71 |
| P1 | 2,256,676 | 564 | 503 (13%) | 0 |
| P2 | 2,733,114 | 683 | 1,947 (49%) | 0 |
| P3 | 2,675,367 | 669 | 1,320 (33%) | 0 |
Every arm wrote all 4,000 posts. Between 13% and 49% of an arm’s posts reached the 1,024-token limit and end mid-text; they are counted as they came, as registered, and the length analysis below shows what that does to the rates. Posts with no words are ones the model ended at once or that hold no word character, and they count as zero words. They cluster in B1, B2, N1 and N3, and the post-trained checkpoints wrote none. The checks made on the text and the counts before the analysis are in docs/collection-log.md.
3 Confirmatory results
3.1 The family
| Contrast | Ratio | Holm p | Verdict | |
|---|---|---|---|---|
| H1 | B1 against the stage-1 sample | 0.43 [0.27, 0.70] | 0.007 | evidence against |
| H2 | Midtraining sample against the stage-1 sample | 0.41 [0.35, 0.47] | < 0.001 | evidence against |
| H3 | B2 against B1 | 6.07 [3.63, 10.14] | < 0.001 | supported |
| H4 | N1 against N2 | 2.82 [2.11, 3.76] | < 0.001 | favours N1 |
| H5 | SFT responses against B3 | 0.06 [0.06, 0.07] | < 0.001 | evidence against |
| H6 | P1 against B3 | 2.20 [1.94, 2.50] | < 0.001 | supported |
| H7 | Chosen against rejected DPO responses | 1.33 [1.20, 1.48] | < 0.001 | supported |
| H8 | P2 against P1 | 1.24 [1.14, 1.36] | < 0.001 | supported |
| H9 | P3 against P2 | 1.14 [1.05, 1.23] | 0.007 | supported |
The end-of-stage-1 checkpoint uses the family less than its training text: B1’s rate is 0.43 [0.27, 0.70] times the stage-1 sample’s, evidence against H1. The midtraining text also carries less of it than the stage-1 text, 0.41 [0.35, 0.47], evidence against H2. Midtraining still raises the model’s rate: B2 is 6.07 [3.63, 10.14] times B1, which supports H3. Of the two midtraining mixes run from the same checkpoint, the one with more web, QA and instruction data and no math, code or reasoning traces gives the higher rate: N1 is 2.82 [2.11, 3.76] times N2, which the registration reads as favouring the QA and instruction data. The mixes differ in several categories at once, and the third of them is web text, which carries the family at a higher rate than either (Table 7), so which category drives the difference is an interpretation and not one this contrast can settle.
The SFT responses use the family at 0.06 [0.06, 0.07] times the final base model’s rate, evidence against H5, yet the SFT checkpoint uses it 2.20 [1.94, 2.50] times as often as B3. The preference data leans the same way as the model. Chosen responses use the family 1.33 [1.20, 1.48] times as often as the rejected responses to the same prompts. DPO then raises the model’s rate by 1.24 [1.14, 1.36] and RLVR by a further 1.14 [1.05, 1.23]. H6 to H9 are supported.
3.2 The composite
| Contrast | Ratio | Holm p | Verdict | |
|---|---|---|---|---|
| H1 | B1 against the stage-1 sample | 0.92 [0.48, 1.77] | 1.000 | not supported |
| H2 | Midtraining sample against the stage-1 sample | 0.55 [0.31, 0.96] | 0.303 | not supported |
| H3 | B2 against B1 | 1.35 [0.66, 2.77] | 1.000 | not supported |
| H4 | N1 against N2 | 1.31 [0.65, 2.63] | 1.000 | not supported |
| H5 | SFT responses against B3 | 0.13 [0.07, 0.24] | < 0.001 | evidence against |
| H6 | P1 against B3 | 1.73 [0.90, 3.32] | 0.649 | not supported |
| H7 | Chosen against rejected DPO responses | 1.31 [0.72, 2.37] | 1.000 | not supported |
| H8 | P2 against P1 | 0.89 [0.47, 1.69] | 1.000 | not supported |
| H9 | P3 against P2 | 1.18 [0.62, 2.24] | 1.000 | not supported |
On the composite only H5 reaches significance, again as evidence against: the SFT responses’ effect is 0.13 [0.07, 0.24] times B3’s. Every other contrast is not supported. The frame-by-source SD is 0.71 on the log scale. Along the training line some frames rise steeply, others hold steady and F04 falls, so an average that weights each of the fifteen equally on the log scale moves little and carries wide intervals. The per-frame rates below show which frames move.
3.3 The family along the training line
| Source | Kind | Words | Hits | Per million |
|---|---|---|---|---|
| Stage-1 sample | training text | 200,443,988 | 4,693 | 23.4 [20.8, 26.1] |
| B1, end of stage 1 | model output | 1,873,609 | 19 | 10.1 [5.4, 14.9] |
| Midtraining sample | training text | 240,035,190 | 2,277 | 9.5 [8.5, 10.5] |
| B2, end of midtraining | model output | 1,836,634 | 113 | 61.5 [48.9, 74.1] |
| B3, final base | model output | 2,408,514 | 477 | 198.0 [177.1, 219.0] |
| SFT responses | training text | 411,272,812 | 5,178 | 12.6 [12.2, 13.0] |
| P1, SFT | model output | 2,256,676 | 984 | 436.0 [406.2, 465.9] |
| DPO chosen responses | training text | 84,142,359 | 2,120 | 25.2 [23.9, 26.5] |
| DPO rejected responses | training text | 77,939,812 | 1,477 | 19.0 [17.1, 20.8] |
| P2, DPO | model output | 2,733,114 | 1,480 | 541.5 [512.0, 571.0] |
| P3, final instruct | model output | 2,675,367 | 1,647 | 615.6 [583.9, 647.3] |
| N1, Gen-QA mix from 2T | model output | 1,891,439 | 217 | 114.7 [97.0, 132.5] |
| N2, math-code-thinking mix from 2T | model output | 2,014,023 | 82 | 40.7 [30.8, 50.6] |
| N3, Round 5 mix from 2T | model output | 1,862,425 | 154 | 82.7 [67.2, 98.2] |
The model’s rate rises at every stage after stage 1: by a factor of 6.07 in midtraining, 3.22 in long-context training (a secondary contrast), 2.20 in SFT, 1.24 in DPO and 1.14 in RLVR. From B1’s 10.1 per million words it reaches 615.6 in the final instruct model. None of the training text comes near the model’s rate once midtraining has run: it runs between 9.5 and 25.2 per million, against 61.5 for B2.
4 Secondary results
Everything in this section is registered as secondary and reported with unadjusted intervals.
| Contrast | Family | Composite |
|---|---|---|
| B3 against B2 | 3.22 [2.56, 4.05] | 1.54 [0.77, 3.05] |
| N3 against N1 | 0.72 [0.57, 0.92] | 1.11 [0.56, 2.21] |
| N3 against N2 | 2.03 [1.49, 2.76] | 1.45 [0.73, 2.91] |
| B2 against the midtraining sample | 6.49 [5.15, 8.17] | 2.28 [1.20, 4.34] |
| P1 against the SFT responses | 34.63 [32.13, 37.33] | 13.55 [7.36, 24.94] |
Long-context training raises the family’s rate by 3.22 [2.56, 4.05], and by 1.54 [0.77, 3.05] on the composite. Its data, two-thirds of it midtraining data, was not counted, so this study cannot say what in that stage raises the rate.
The Round 5 mix from 2T lands between the other two: N3 is 0.72 [0.57, 0.92] times N1 and 2.03 [1.49, 2.76] times N2.
The two checkpoints set against the text of their own stage both outrun it. B2 is 6.49 [5.15, 8.17] times the midtraining sample and P1 is 34.63 [32.13, 37.33] times the SFT responses, and both hold on the composite.
Adding the partial F10 to the family leaves every confirmatory ratio on the same side of 1, with an interval that excludes 1. Where F10 makes up much of the hits and changes little between the two sources, it pulls the ratio towards 1: H3’s falls from 6.07 to 3.20.
The final base and final instruct checkpoints are the ones Study 2 ran at the same settings, with new seeds, and their family rates reproduce Study 2’s. B3 is 1.07 [0.93, 1.25] times Study 2’s base-1.0, and P3 is 1.02 [0.95, 1.10] times its instruct-1.0.
4.1 Where the family sits in the training text
| Category | Files | Words | Hits | Per million |
|---|---|---|---|---|
| Web pages | 275 | 65.8 million | 1,488 | 22.6 [20.5, 24.7] |
| QA (synth) | 160 | 33.4 million | 398 | 11.9 [8.1, 15.7] |
| Instruction (synth) | 60 | 14.7 million | 120 | 8.2 [4.7, 11.7] |
| Thinking (synth) | 87 | 20.0 million | 82 | 4.1 [2.5, 5.7] |
| Math (synth) | 188 | 46.1 million | 167 | 3.6 [2.7, 4.6] |
| PDFs | 51 | 12.0 million | 8 | 0.7 [0.2, 1.2] |
| Code | 97 | 24.0 million | 14 | 0.6 [0.2, 1.0] |
| Python (synth) | 97 | 24.0 million | 0 | no hits |
In the midtraining sample, 65% of the family’s hits come from web text, which runs at 22.6 per million, close to the stage-1 sample’s rate. The synthetic QA sources run at 11.9 as a group and the instruction data at 8.2, with one QA source, Nemotron Synth QA, the highest of the 24 at 32.6. Code, mathematics, reasoning traces and PDFs carry little. A mix that trades mathematics, code and reasoning traces for web, QA and instruction data, as the Gen-QA mix does, therefore holds more of the family per word, which may account for part of N1’s higher rate. Web text is the largest part of that: it is the highest-rate category here and supplies most of the family’s hits, so the web share N1 raises is at least as good a candidate as the QA and instruction data. The two experimental mixes’ own text was not measured, and the released mix, which was, differs from both.
| Dataset | Conversations | Words | Hits | Per million |
|---|---|---|---|---|
| Wildchat | 302,406 | 108,185,379 | 3,836 | 35.5 [34.2, 36.7] |
| Dolci Instruct Precise IF | 136,833 | 26,028,910 | 784 | 30.1 [27.7, 32.6] |
| WildGuardMix | 49,373 | 3,621,643 | 91 | 25.1 [19.9, 30.4] |
| WildJailbreak | 49,965 | 10,374,813 | 197 | 19.0 [16.2, 21.8] |
| CoCoNot | 10,957 | 1,925,451 | 10 | 5.2 [2.0, 8.4] |
| Evol CodeAlpaca | 107,270 | 24,886,685 | 98 | 3.9 [3.1, 4.8] |
| Dolci Instruct OpenThoughts3+ Science | 99,268 | 38,568,362 | 129 | 3.3 [2.7, 4.0] |
| OpenAssistant | 7,132 | 1,407,936 | 2 | 1.4 [0.0, 3.4] |
| FLAN | 89,981 | 1,720,455 | 2 | 1.2 [0.0, 2.8] |
| Dolci Instruct Tool Use | 227,579 | 27,520,108 | 19 | 0.7 [0.4, 1.0] |
| Aya | 99,987 | 11,734,073 | 2 | 0.2 [0.0, 0.4] |
| Verifiable Reasoning | 310,572 | 35,004,940 | 4 | 0.1 [0.0, 0.2] |
| Dolci Instruct Python Algorithms | 186,345 | 19,124,718 | 1 | 0.1 [0.0, 0.2] |
| Tulu 3 Persona MATH | 149,958 | 69,789,571 | 3 | 0.0 [0.0, 0.1] |
| Hardcoded Data | 69 | 4,475 | 0 | no hits |
| Logic Puzzles | 159,882 | 6,488,542 | 0 | no hits |
| OpenMathInstruct 2 | 50,000 | 4,912,582 | 0 | no hits |
| SciRiff | 4,557 | 201,001 | 0 | no hits |
| TableGPT | 5,000 | 165,902 | 0 | no hits |
| Tulu 3 Persona Algebra | 19,999 | 7,475,146 | 0 | no hits |
| Tulu 3 Persona GSM | 49,980 | 9,636,276 | 0 | no hits |
| Tulu 3 Persona Python | 34,999 | 2,495,844 | 0 | no hits |
In the SFT set the family comes mostly from one dataset: the WildChat prompts with responses written by GPT-4.1 hold 26% of the words and 74% of the family’s hits, at 35.5 per million. That is 8% of the SFT checkpoint’s rate. The mathematics, code and puzzle datasets carry little or none.
In the DPO set, qwen3-no_reasoning-32b wrote 52% of the chosen responses, at 22.9 per million, and qwen3-no_reasoning-0.6b wrote 57% of the rejected ones, at 18.2. H7’s difference therefore reflects which models wrote each side as well as which responses were preferred. The per-model rates are in data/study4/results-dpo-models.csv.
5 Exploratory results
Everything in this section is registered as exploratory. It is descriptive and uncorrected.
5.1 Frame by frame
| Frame | Stage 1 | B1 | B2 | B3 | P1 | P2 | P3 |
|---|---|---|---|---|---|---|---|
| F01 | 1.3 | 2.1 | 1.6 | 3.7 | 13.3 | 19.0 | 20.9 |
| F02 | 0.4 | 0.0 | 0.5 | 1.2 | 2.2 | 2.2 | 3.0 |
| F03 | 8.8 | 12.3 | 10.9 | 8.3 | 66.9 | 72.4 | 57.6 |
| F04 | 4.1 | 3.7 | 3.3 | 1.2 | 0.4 | 0.0 | 0.7 |
| F05 | 0.4 | 0.0 | 0.0 | 0.4 | 0.0 | 0.0 | 0.4 |
| F06 | 0.5 | 0.0 | 0.5 | 0.4 | 4.4 | 0.4 | 0.7 |
| F07 | 10.2 | 5.3 | 26.7 | 61.0 | 82.4 | 81.6 | 101.3 |
| F08 | 2.6 | 0.5 | 4.9 | 32.0 | 8.9 | 14.3 | 16.4 |
| F09 | 10.1 | 4.3 | 29.4 | 104.6 | 340.3 | 445.3 | 497.1 |
| F10 | 15.1 | 13.3 | 13.6 | 12.5 | 43.0 | 34.8 | 28.8 |
| F11 | 0.2 | 0.5 | 1.1 | 0.4 | 0.0 | 0.4 | 0.0 |
| F12 | 0.0 | 1.1 | 0.0 | 0.0 | 0.0 | 0.0 | 0.0 |
| F13 | 41.3 | 38.4 | 40.8 | 34.5 | 43.9 | 32.9 | 47.1 |
| F14 | 19.8 | 19.7 | 23.4 | 17.4 | 21.3 | 17.2 | 14.2 |
| F15 | 0.7 | 3.2 | 0.0 | 2.5 | 3.1 | 2.6 | 5.2 |
F09, is n't just, carries most of the family’s rise. It runs at 10.1 per million in the stage-1 sample, at 104.6 in the final base model and at 497.1 in the final instruct model, where it makes up 81% of the family’s hits. F07, it 's not just, rises at each step from B1 to B3, and F08, not just * but, peaks in B3. F06, the full it 's not * , it 's, is rare everywhere, and F10 holds near the stage-1 rate in every base-model checkpoint and runs higher in the post-trained ones.
The other families move differently. Two announce-the-point frames rise mainly at SFT: F01, here 's the thing, from 3.7 per million in B3 to 20.9 in P3, and F03, here 's what, from 8.3 to 57.6. F04, the thing is, falls from 4.1 in the stage-1 sample to 0.7 in P3, while F13, and that 's, and F14, turns out, stay near their stage-1 rates throughout. That spread is what the composite’s frame-by-source SD records.
5.2 Length and truncation
| Arm | Finished posts | Per million, finished | Cut-off posts | Per million, cut off |
|---|---|---|---|---|
| B1 | 2,404 | 14.1 | 1,596 | 8.1 |
| B2 | 2,616 | 57.8 | 1,384 | 64.1 |
| B3 | 2,521 | 214.6 | 1,479 | 180.7 |
| N1 | 2,604 | 136.2 | 1,396 | 99.0 |
| N2 | 2,348 | 26.1 | 1,652 | 49.7 |
| N3 | 2,586 | 94.8 | 1,414 | 74.1 |
| P1 | 3,497 | 460.3 | 503 | 304.4 |
| P2 | 2,053 | 598.8 | 1,947 | 489.8 |
| P3 | 2,680 | 678.6 | 1,320 | 508.8 |
In seven of the nine arms, posts that finished carry the family at a higher rate than posts cut at the token limit, so truncation lowers those arms’ rates, more so where more posts were cut. The order along the training line holds within each kind of post: among finished posts and among cut-off posts alike, each of B1, B2, B3, P1, P2 and P3 has a higher rate than the one before.
6 What this does not establish
- One model. Olmo 3 7B was chosen because every stage’s checkpoint and data are published. Nothing here says how other models acquire the frames.
- One prompt and one genre. Every text answers
Write a blog post titled:followed by a title from a 2012 blog. - Genre in the comparisons with training text. H1, H2, H5 and the secondary comparisons of a checkpoint with its own data set prompted blog posts against training text of every kind, from code to encyclopedia articles. The comparisons between checkpoints hold the prompts fixed. Within the training text, the sources nearest to blog writing, web text and GPT-4.1’s WildChat responses, still run far below the post-trained checkpoints.
- What a stage bundles. A difference between successive checkpoints is the effect of all the training between them: the data, the learning-rate schedule and, at the step from base to SFT, the prompt format and the chat template’s default system prompt.
- One experiment on the mixes. N1 to N3 start from a stage-1 checkpoint at 2 trillion tokens, earlier than the one the released model continued from, and their mixes differ in several categories at once.
- Stages measured on one side. The long-context stage’s data was not counted, and the RLVR data holds no responses. For both, only the change in the model’s output is measured.
- Few hits at the start. B1’s rate rests on 19 hits, so its interval is wide.
- A midtraining sample by token share. The sources are allocated by their shares of tokens and rates are per word, so the pooled sample matches the mix’s composition only approximately, and the olmOCR source’s rate describes its unredacted documents only (S4-D4).
- Truncated texts. Between 13% and 49% of an arm’s posts were cut at 1,024 tokens. The length analysis above shows they lower the rates without changing the order.
- Fixed strings chosen as typical of model output. As in Studies 1 to 3, these are rates per million of fifteen frozen strings, chosen because they seem typical of machine-generated prose. Density, sentence position and paraphrases are not measured.
- Text that cannot be regenerated exactly. Each generation’s seed fixes its sampling request, but vLLM’s arithmetic varies with the batch a request shares, so a rerun gives different text.
data/study4/generated-sha256.txtidentifies the files that were counted.
7 Reproducing this
The repository holds the generation and counting code (python/generate_study4.py, python/count_study4.py), the protocol with every checkpoint’s and dataset’s pinned commit (docs/study4-protocol.md), the counts (data/study4/counts.csv, with one row per unit of text in data/study4/units-*), one record per midtraining file read (data/study4/midtraining/), and the record of the run (docs/collection-log.md).
Rscript R/analyse-study4.R recomputes every number in this document from those files, and python python/count_study4.py --collect data/study4 rebuilds the tables from the generated text and the per-shard counts, which git does not hold.