Studies 2 and 3: rhetorical frames in Olmo 3’s prose and its training data

englishRegisterStudy — results from the pre-registered analyses

Author

Cesaire Tobias

Published

29 September, 2026

1 What this measures

Study 1 found that American web English in 2012 used fifteen frozen rhetorical frames more often than the other four national varieties measured, which used them at between 0.50 and 0.69 times the American rate. The frames were chosen because they seem typical of machine-generated prose: announcing a point before making it, the it’s not X, it’s Y contrast, the epigrammatic assertion. The argument behind the whole study is that such prose reads as foreign to readers of other varieties partly because it reads as American.

Two further studies test the next steps. Each was registered before its own data existed.

  • Study 2 (https://osf.io/qjgtc) asks whether a language model’s unprompted prose sits at the American end of Study 1’s spread. H1 predicts that the model’s composite rate over the fourteen frames Study 1 could count is higher than each of GB, IE, AU and ZA. The comparison against US is reported without a predicted direction.
  • Study 3 (https://osf.io/ngt3m) asks whether the model uses the frames more often than the text it was trained on. A model that reproduces its training data’s register is working correctly, so H1 predicts amplification: a composite rate above that of the web portion of the pretraining corpus.

The model is Olmo 3 7B, chosen because Ai2 published the corpus it was trained on, which is what makes Study 3 possible. Four conditions were generated: the base and instruction-tuned models (allenai/Olmo-3-1025-7B, allenai/Olmo-3-7B-Instruct), each at unmodified sampling (temperature 1.0, top-p 1.0) and at typical deployment settings (0.7 and 0.9). The confirmatory condition in both studies is the base model at unmodified sampling, base-1.0. Each condition wrote one post for each of 4,000 titles drawn from GloWbE’s blog pages (Davies, 2013), 800 from each variety, prompted with Write a blog post titled: and the title, and nothing else.

Four deviations were declared before the data they affect existed, and each was filed as a registration update; all four are in docs/deviations.md. Two concern the generated text. The instruct model receives the default system prompt its chat template inserts (S2-D1), and every condition covers all 4,000 titles even where that takes it past the registered two million words (S2-D2). The other two concern the Dolma draw (S3-D1 and S3-D2).

2 The generated text

Table 1: The four conditions as generated on 28 September 2026.
Condition Generations Words Reached 1,024 tokens Empty
base-1.0 4,000 2,426,034 1,503 (38%) 2
base-0.7 4,000 2,327,984 1,050 (26%) 0
instruct-1.0 4,000 2,680,977 1,284 (32%) 0
instruct-0.7 4,000 2,614,894 918 (23%) 0

Each condition passed two million words on its first pass over the titles, so no title was reused and topic is the same across all four. Between 23% and 38% of generations reached the 1,024-token limit and end mid-text. They are counted as they came, as registered, and the length analysis below shows what that does to the rates. The two empty generations are posts the base model ended before writing anything. The checks made on the text before any frame was counted are in docs/collection-log.md.

3 Study 2: the model against the national varieties

3.1 The confirmatory test

Table 2: Confirmatory result: the composite rate of the 14 shared frames in base-1.0, relative to each variety, with unadjusted 95% intervals.
Variety base-1.0 against the variety Holm p Supports H1
US 1.19 [0.60, 2.38] 0.587 no direction predicted
GB 1.76 [0.88, 3.51] 0.226 no
IE 2.25 [1.12, 4.49] 0.102 no
AU 1.86 [0.93, 3.71] 0.226 no
ZA 2.69 [1.33, 5.41] 0.047 yes

H1 is supported against ZA only, where the model’s rate is 2.69 times South Africa’s with a Holm-adjusted p of 0.047. Against GB, IE and AU the point estimates are above 1 and the adjusted p-values are not below 0.05. The intervals in the table are unadjusted, which is why Ireland’s excludes 1 while its adjusted p is 0.102. Against American English the ratio is 1.19, with an interval that runs from well below the American rate to more than twice it.

The intervals are wide because the frames disagree about the model. The frame-by-source SD is 0.78 on the log scale, against 0.15 among the varieties in Study 1, so the model departs from the varieties frame by frame far more than the varieties depart from each other. The exploratory section below shows where.

Table 3: The 14 shared frames summed, for each variety and each condition.
Source Hits Per million 95% interval
US 63,080 163.1 [161.8, 164.4]
GB 47,810 123.3 [122.2, 124.5]
IE 9,467 93.7 [91.8, 95.6]
AU 18,044 121.7 [120.0, 123.5]
ZA 3,991 88.0 [85.3, 90.7]
base-1.0 644 265.5 [245.3, 286.8]
base-0.7 453 194.6 [177.1, 213.4]
instruct-1.0 2,119 790.4 [757.1, 824.8]
instruct-0.7 2,292 876.5 [841.0, 913.2]

Summed, base-1.0’s rate is 1.63 times the American one. The model’s ratio of 1.19 is smaller because it averages over frames on the log scale, where a large excess on a few frames counts for less than it does in a sum.

3.2 The four conditions

Table 4: The registered comparisons among the conditions, Holm-corrected within the set of three.
Comparison Ratio Holm p
base-1.0 against base-0.7 1.57 [1.00, 2.49] 0.052
base-1.0 against instruct-0.7 0.52 [0.34, 0.81] 0.014
base-1.0 against instruct-1.0 0.50 [0.32, 0.76] 0.013

These comparisons use 12 frames: F05, F11, F12 have no hits in any condition and leave the model, as registered. The instruction-tuned model uses the frames at about twice the base model’s rate, at either setting. This matches Reinhart et al. (2025), who compared Llama 3 and GPT-4o with human writing on Biber’s grammatical and rhetorical features and found the differences larger for instruction-tuned models than for base models. The base model at deployment settings sits below the same model at unmodified sampling: base-1.0 is 1.57 [1.00, 2.49] times base-0.7, with an adjusted p of 0.052. Both instruct conditions include the default system prompt, so the difference between base and instruct includes whatever that prompt contributes.

4 Study 3: the model against its training data

4.1 The primary test

The input is a 200-million-word sample of common_crawl, the web portion of Dolma 3’s stage-1 mix and 76% of its tokens by Ai2’s published composition, drawn as registered and recorded in docs/collection-log.md.

Table 5: Primary result, one uncorrected test, and its registered robustness check.
Frames base-1.0 against common_crawl p
All fifteen 1.72 [0.65, 4.53] 0.250
Without F10 and F13 1.94 [0.62, 6.08] 0.232

H1 is not supported. The confirmatory condition’s rate is 1.72 times the input’s, with an interval from 0.65 to 4.53 and p = 0.250. The data are consistent both with the model reproducing its input’s rate and with it exceeding that rate several times over. Removing the two partial frames leaves the conclusion where it is. The frame-by-source SD is 1.16, larger than Study 2’s, and again it is the frames’ disagreement that makes the interval wide.

4.2 The other conditions

Table 6: Secondary results, Holm-corrected within the set of three.
Condition Against common_crawl Holm p
base-0.7 0.93 [0.34, 2.54] 0.876
instruct-0.7 2.97 [1.14, 7.72] 0.075
instruct-1.0 3.06 [1.18, 7.97] 0.075

Both instruct conditions sit about three times above the input, with adjusted p-values of 0.075. The base model at deployment settings is close to the input’s rate.

4.3 The whole mix

The registered comparison against the whole stage-1 mix gives 2.81 [2.60, 3.04]. It rests on an approximation the registration names, token shares applied to per-word rates, and it measures something different from the primary test. It is a ratio of summed rates, base-1.0’s 266.7 per million against a token-weighted mix rate of 94.7, and its interval reflects counting noise alone. Frame disagreement, which widens every other interval in this document, does not enter it, and a sum is dominated by the frames with the most hits, which here are the contrastive frames the model over-uses. The mix rate is also pulled down by the code, mathematics, arXiv, science PDF and Wikipedia subsets, which use the frames far less than web text does.

Table 7: Each subset’s summed rate over all fifteen frames, and the confirmatory condition against it, with unadjusted intervals.
Subset Per million base-1.0 against the subset
common_crawl 115.4 1.72 [0.65, 4.53]
finemath-3plus 59.8 3.05 [1.14, 8.17]
rpj-proofpile-arxiv 39.7 12.88 [4.57, 36.34]
olmocr_science_pdfs 32.8 9.92 [3.59, 27.40]
dolma1_7-wiki-en 16.4 31.40 [10.69, 92.19]
stack_edu 9.2 39.51 [13.25, 117.85]

The olmOCR draw reached 15 of the subset’s 21 topics. The ones it missed hold 28% of its compressed bytes, health and education most of that. Reweighting the per-topic rates to byte shares over the topics reached moves olmOCR’s rate from 32.8 to 27.4 per million, and the whole-mix ratio from 2.81 to 2.84. The olmOCR rate describes the unredacted documents only (S3-D2).

5 What drives the composite

Everything in this section is registered as exploratory. It is descriptive and uncorrected, and no conclusion rests on a single frame.

5.1 By family

Table 8: Each family’s summed rate over the shared frames, relative to American English in 2012.
Source announce contrastive epigram reveal
GB 0.64 0.86 0.79 0.63
IE 0.54 0.60 0.62 0.46
AU 0.65 0.79 0.80 0.64
ZA 0.54 0.61 0.54 0.43
common_crawl 0.68 0.84 0.65 0.64
base-1.0 1.03 4.33 0.45 0.54
base-0.7 0.64 3.05 0.56 0.18
instruct-1.0 3.87 14.06 0.79 0.43
instruct-0.7 3.94 15.95 0.85 0.27

The model differs from the varieties in shape as well as level. On the announce-the-point frames, which separated the varieties most in Study 1, base-1.0 is at 1.03 times the American rate. On the epigrammatic frames it is at 0.45, below every variety, and on the single reveal frame at 0.54, among the lower varieties. On the contrastive frames it is at 4.3 times the American rate, and the instruct model at 14 times. The contrastive family separated the varieties least in Study 1, where British or Australian writers used two of its four countable frames more often than Americans did.

The model’s input, common_crawl, sits below American English of 2012 on every family, near the other four varieties. Its rates rest on a different word count from Study 1’s, and it is a filtered crawl of English from everywhere, so the comparison is loose. It does show that the stage-1 web text the model was trained on has no contrastive excess of its own.

5.2 Frame by frame

Table 9: The shared frames, ordered by how far base-1.0 exceeds its input. The first column is the US rate over the mean of the other four varieties.
Frame Query US against the others, Study 1 base-1.0 against common_crawl instruct-1.0 against base-1.0
F09 is n't just 1.69 10.22 4.70
F08 not just * but 1.05 8.66 0.60
F07 it 's not just 1.23 5.60 1.82
F02 here 's the * part 2.57 5.58 2.35
F01 here 's the thing 2.18 3.24 5.16
F15 the real question is 2.27 1.82 3.02
F03 here 's what 2.10 1.69 3.82
F10 not * , but * 1.45 0.90 2.58
F14 turns out 1.85 0.85 0.79
F13 and that 's 1.45 0.69 1.77
F04 the thing is 1.11 0.30 1.21
F05 what 's interesting is 2.23 0.00 no hits
F11 that 's the whole point 1.82 0.00 no hits
F12 which is the point 1.64 0.00 no hits

The registration asks whether the frames the model most over-produces are the ones Study 1 found most American. Across the 14 frames, the rank correlation between the first two ratio columns is -0.05. The three frames the model over-produces most against its input, F09, F08, F07, are all contrastive. None is among Study 1’s five most American frames, and F08 is its least American. The three frames the base model never produced, F05, F11, F12, all ran at more than 1.6 times the other varieties’ rate in American English.

Post-training raises 9 of the 11 frames the base model produced, F01 most of all.

5.3 How widely the hits spread

A base model at unmodified sampling can fall into loops, and a few looping texts would produce a large count. The contrastive hits are spread thinly instead. F09 appears in 220 of base-1.0’s 4,000 texts and never more than 3 times in one, and in 1,049 of instruct-1.0’s, at most 4 in one.

5.4 Length and truncation

Table 10: Family rates per million over the shared frames, by whether the text finished or was cut off.
Condition Ending Texts Mean words Announce Contrastive Epigram Reveal
base-1.0 finished 2,497 493 17.9 207.4 27.6 14.6
base-1.0 cut off 1,503 796 29.3 185.6 29.3 19.2
base-0.7 finished 2,950 498 12.2 136.8 36.7 5.4
base-0.7 cut off 1,050 818 18.6 142.1 33.8 5.8
instruct-1.0 finished 2,716 631 82.8 738.9 59.5 15.7
instruct-1.0 cut off 1,284 753 97.3 460.5 34.2 9.3
instruct-0.7 finished 3,082 622 85.0 809.4 63.6 8.9
instruct-0.7 cut off 918 760 101.8 491.8 27.2 7.2

Texts that reached the token limit were cut before their ending. If the frames cluster at the ends of posts, truncation lowers a condition’s rate. For the base model it makes little difference: base-1.0’s epigrammatic rate is 27.6 per million in texts that finished and 29.3 in texts that were cut, so its low epigrammatic rate is not an effect of truncation. The instruct model’s finished texts carry far more contrastive and epigrammatic frames than its truncated ones. Truncation therefore lowers the instruct rates, and the instruct conditions’ excess is if anything understated.

5.5 Variation between the drawn files

The 200 common_crawl files differ widely from one another: the middle 90% run from 52.8 to 214.0 per million over all fifteen frames. No single file matters much. Dropping any one of them moves the subset’s rate by at most 1.0%.

6 What this does not establish

  • One model. Olmo 3 7B was chosen because its training corpus is published. Nothing here says how GPT, Claude or Gemini write, and for them the input side cannot be measured at all.
  • One prompt and one genre. Every text answers Write a blog post titled: followed by a title from a 2012 blog. Other prompts and genres may draw on other registers.
  • Twelve years. Study 1’s baseline is web English collected in December 2012, and the model card gives December 2024 as the cutoff of the model’s data. A gap between the model and the varieties mixes register with whatever changed in web writing in between. Study 3’s input comes from the model’s own period.
  • Different denominators. Study 1’s rates divide by GloWbE’s own word counts, and Studies 2 and 3 divide by this study’s word rule (docs/frame-mapping.md). Comparisons within Studies 2 and 3 share one footing. Comparisons back to Study 1 do not, and the direction of any bias is unknown.
  • Fourteen frames against Study 1. F06 is counted in Studies 2 and 3, but a free GloWbE account could not search it.
  • A system prompt in the instruct conditions. The template inserts a default system prompt (S2-D1), so what post-training adds cannot be separated from what that prompt adds.
  • Stage 1 only. The input sample comes from the stage-1 mix, 5.93 of the 6.08 trillion tokens the base model’s card lists across its three pretraining stages. The two later stages were not sampled, and they differ in kind: Ai2’s card for the midtraining mix labels most of it synthetic, including instruction data and reasoning traces from other models, among them QwQ, Gemini and Llama Nemotron. The instruct model’s post-training adds supervised examples that include responses written by GPT-4.1, and preference pairs partly judged by a GPT model. The base model’s departure from its stage-1 input may therefore come partly from model-written text it saw later, and neither that departure nor the instruct model’s excess can be traced to a stage.
  • An incomplete olmOCR draw. It missed six of the subset’s topics and cannot see redacted documents. Both affect the whole-mix figure only.
  • Truncated texts. Between 23% and 38% of generations were cut at 1,024 tokens. The length analysis above bounds what that does.
  • Fixed strings. As in Study 1, these are rates per million. Density, sentence position and near-obligatory use are not measured.
  • Frames chosen as typical of model output. A high model rate on some of them is what got them chosen. The studies ask where the model’s rate sits against the varieties and against its input, which the choice does not settle, but a different fifteen frames could tell a different story.
  • Text that cannot be regenerated exactly. Each generation’s seed fixes its sampling request, but vLLM’s arithmetic varies with the batch a request shares, so a rerun gives different text. data/generated-sha256.txt identifies the files that were counted.

7 Reproducing this

The repository holds the generation code (python/generate.py), the titles (data/topics.csv), the counts of the generated text (data/counts-generated.csv, with each text’s own counts in data/counts-generated-by-text.csv and the run’s totals in data/generated-run.csv), the Dolma draw’s per-file records (data/dolma/) and their counts (data/counts-dolma.csv and data/counts-dolma-topic.csv), and the record of both runs (docs/collection-log.md).

Rscript R/analyse.R, Rscript R/analyse-generated.R and Rscript R/analyse-training.R recompute every number in this document from those files. python python/count_generated.py data/generated recounts the generated text into the tables above.

8 References

Davies, M. (2013) “Corpus of Global Web-Based English: 1.9 billion words from speakers in 20 countries (GloWbE).” Available at: https://www.english-corpora.org/glowbe/.
Reinhart, A. et al. (2025) “Do LLMs write like humans? Variation in grammatical and rhetorical styles,” Proceedings of the National Academy of Sciences, 122(8), p. e2422455122. Available at: https://doi.org/10.1073/pnas.2422455122.