Study 1: rhetorical frames across national varieties of web English

englishRegisterStudy — results from the pre-registered analysis

Author

Cesaire Tobias

Published

25 September, 2026

1 What this measures

Fifteen rhetorical frames were frozen before any querying, chosen from an impression of what language models overproduce: announcing a point before making it, the it’s not X, it’s Y contrast, the epigrammatic assertion. The question is whether they are features of American web register rather than of English generally.

The instrument is GloWbE (Davies, 2013), 1.9 billion words of web text from twenty countries, collected in December 2012. Five national components are used: US, GB, IE, AU and ZA. The collection date is a design choice rather than an accident. Any corpus collected after 2022 contains generated text in unknown proportion, so a human baseline drawn from one would be circular.

The design, the frames and the analysis were registered at https://osf.io/48wjn on 21 September 2026, before any count was retrieved. Two deviations were declared during collection and filed as registration updates; both are in docs/deviations.md.

2 What the interface allowed

Table 1: The free tier’s limits decided the shape of the analysis.
What the interface allowed Frames
Counted only with sections combined 6
Counted in each section 8
Not countable on a free account 1

A free english-corpora.org account refuses a section-restricted search when every slot in the query occurs 1,500,000 times or more, and caps a search string at five tokens. The first blocks six frames, which is the whole contrastive family apart from one, plus two others. The second blocks F06, it 's not * , it 's, entirely: it is seven tokens.

The consequence is deviation D2. The confirmatory analysis covers the 14 countable frames with both sections combined, and the registered per-section model becomes a secondary analysis on the 8 frames that split. The genre control is what this costs: it now covers 8 frames rather than fifteen, and not the frames whose rates a genre explanation would most plausibly account for.

3 Results

3.1 The composite

Table 2: Confirmatory result: the composite rate of the 14 countable frames in each variety, relative to American, with 95% intervals.
Variety Ratio against US Holm p Supports H1
GB 0.69 [0.61, 0.78] < 0.0001 yes
IE 0.54 [0.47, 0.62] < 0.0001 yes
AU 0.66 [0.58, 0.76] < 0.0001 yes
ZA 0.50 [0.43, 0.57] < 0.0001 yes

Every variety uses the frames less often than American English does, and every comparison survives Holm’s correction. The hypothesis was directional and pre-registered, and all four varieties support it.

The frame-by-variety random effect has an SD of 0.149 on the log scale, so frames disagree modestly about how large the gap is. That disagreement is carried into the intervals above, which is what the model is for.

Table 3: The same thing descriptively: all countable frames summed.
Variety Hits Per million 95% interval
US 63,080 163.1 [161.8, 164.4]
GB 47,810 123.3 [122.2, 124.5]
IE 9,467 93.7 [91.8, 95.6]
AU 18,044 121.7 [120.0, 123.5]
ZA 3,991 88.0 [85.3, 90.7]

3.2 Robustness

Two frames are marked partial in data/frames.csv because they match text outside the construction they stand for: not * , but * and and that 's. The second is the larger worry, being roughly ten times the size of any other frame in raw counts. Removing both and refitting moves every ratio slightly further from 1 rather than towards it, so the result does not rest on what those two patterns are really counting.

Table 4: The primary result with and without F10 and F13.
Variety All countable frames Without the partial frames
GB 0.69 [0.61, 0.78] 0.68 [0.58, 0.79]
IE 0.54 [0.47, 0.62] 0.53 [0.45, 0.62]
AU 0.66 [0.58, 0.76] 0.64 [0.55, 0.75]
ZA 0.50 [0.43, 0.57] 0.47 [0.39, 0.56]

3.3 The genre control

Table 5: Secondary analysis: the registered per-section model on the 8 frames the interface splits.
Variety Ratio against US Holm p
GB 0.59 [0.50, 0.70] < 0.0001
IE 0.50 [0.42, 0.60] < 0.0001
AU 0.59 [0.50, 0.70] < 0.0001
ZA 0.47 [0.39, 0.58] < 0.0001

The variety × section interaction test finds no evidence that the gap differs between the corpus’s blog and general sections: likelihood ratio 4.50 on 4 degrees of freedom, p = 0.3423. The registration makes that test the rule for which set is confirmatory, and at a p of 0.05 or above it is the ratios common to both sections, which is what Table 5 reports. That is the weak form of evidence against a pure genre explanation: if the difference between varieties were really a difference between blogs and broadsheets, holding the section constant should have shrunk it, and it does not.

The control is weaker than it looks, in two ways. The corpus’s own composition table heads its General column “General (may also include blogs)”, so the sections are not clean genres; Murphy (2025) abandoned balancing a GloWbE sample across them on the strength of Biber, Egbert and Davies (2015), a paper this study could not obtain and therefore cites at second hand. And the control covers 8 frames, not all of them.

3.4 Where the pattern is not uniform

Table 6: Every per-frame cell where a variety exceeds American.
Frame Query Variety Ratio
F04 the thing is GB 1.03 [0.97, 1.08]
F07 it ’s not just GB 1.08 [1.04, 1.12]
F08 not just * but AU 1.03 [0.89, 1.18]
F08 not just * but GB 1.18 [1.07, 1.30]

52 of the 56 per-frame ratios are below 1. The exceptions are British and Australian, and mostly contrastive frames. By family, the median ratio runs from 0.47 for the announce frames to 0.73 for the contrastive frames: the announce-the-point moves separate the varieties most, and the contrastive ones least. These ratios are descriptive, uncorrected and registered as such, so this is a lead for the write-up rather than a finding.

4 What this does not establish

  • It is 2012 web English. The register in question is claimed to have spread through Medium, Substack and LinkedIn, which mostly postdate the corpus. Nothing here speaks to English since.
  • It counts fixed strings. The sharper version of the claim is about density and near-obligatory use, and a rate per million is a weak proxy for it. Sentence position, one-sentence paragraphs and closing epigrams are not reachable through the web interface at all.
  • The frames were chosen from intuition. Freezing and registering them controls for choosing after seeing results. It cannot show they were the right fifteen.
  • Country assignment is a proxy. GloWbE assigns pages by site rather than by author nationality. From a spelling test, Murphy (2025) estimates 10–15% of writers in the GB and US components are non-nationals, which biases a variety comparison toward the null: the gaps above are more likely understated than manufactured.
  • No generated text appears anywhere in this study. That American web English used these frames more in 2012 says nothing on its own about why generated prose reads as it does. Study 2 measures model output and Study 3 measures the training data behind it; only with those does the argument about machine writing become testable.

5 Reproducing this

Every count came from the GloWbE web interface with a free account. The repository holds the frozen frames (data/frames.csv), the counts (data/counts.csv and data/counts-combined.csv), the section and component word counts (data/corpus-sizes.csv), and what each query session did, including the refusals (docs/collection-log.md).

Rscript R/analyse.R recomputes every number in this document. Rscript R/check-corpus-sizes.R recomputes the word counts from the corpus’s own metadata download. Rscript R/calibrate-composite.R reruns the simulations behind the choice of model and interval.

6 References

Biber, D., Egbert, J. and Davies, M. (2015) “Exploring the composition of the searchable web: A corpus-based taxonomy of web registers,” Corpora, 10(1), pp. 11–45. Available at: https://doi.org/10.3366/cor.2015.0065.
Davies, M. (2013) “Corpus of Global Web-Based English: 1.9 billion words from speakers in 20 countries (GloWbE).” Available at: https://www.english-corpora.org/glowbe/.
Murphy, M.L. (2025) “Separated by a common im/politeness marker: Please in American and British web-based English,” English Language & Linguistics, 29(4), pp. 781–804. Available at: https://doi.org/10.1017/S1360674324000455.