What 6,500 documents say about a field
A corpus study of 49 ILIAD authors: 2,743 long-form pieces (2003–2026) and 3,653 social posts (Aug 2025–Aug 2026). Everything below is measured, not vibes — though the opinions at the end are mine.
One of three essays from the same study; the companion pieces read the social corpus and the long-form archive separately and in more depth — this is the synthesis.
1. They are not the people they think they are about calibration
Start with the finding that will annoy the most people, because it is the most robust.
Three readers went through three independent 244-post samples of the community’s social output. The phrase “epistemic status” appeared once, once, and zero times. Explicit numeric probabilities about the world: roughly 2–4% of posts. In the long-form archive, explicit epistemic-status headers run 2–9% depending on the bundle.
Meanwhile “I think” is in a fifth of everything, “seems” in a seventh, and 22% of posts contain a parenthetical aside of sixty characters or more, doing the actual hedging work.
These authors did not abandon calibration. They conjugated it. The tell is Paul Rapoport answering a three-part question entirely in graded negations — weak no, no I think, wouldn’t be surprised if it were a weak no — which is not a hedge wrapped around a claim but a claim built out of hedges.
I think this is fine, actually, and I’ll defend it below. But the self-image needs updating: this is a verbally Bayesian culture, not a numerically Bayesian one, and critics who mock the P(doom) thing are mostly mocking a stereotype that the community’s own output does not support.
2. Agent foundations really did decline. Here are the numbers.
Share of long-form pieces touching each theme, era-banded:
| 2008–13 | 2014–18 | 2019–22 | 2023–26 | |
|---|---|---|---|---|
| Agent Foundations & Decision Theory | 18.4% | 25.5% | 24.9% | 13.3% |
| Deep Learning / LLMs | 9.4% | 13.4% | 32.4% | 43.4% |
| Mechanistic Interpretability | 1.6% | 1.7% | 7.1% | 20.0% |
| Evaluation, Oversight & Control | 1.9% | 4.2% | 5.6% | 17.2% |
| Singular Learning Theory | 6.1% | 2.9% | 3.3% | 12.6% |
| Rationality & Epistemics | 49.5% | 34.9% | 36.0% | 30.5% |
| Research Practice & Community | 25.6% | 21.8% | 21.4% | 26.1% |
Interp is up twelvefold. Evals ninefold. Agent foundations fell 48% from its peak. Rationality has been in continuous decline since 2008 and is now smaller than deep learning by thirteen points.
SLT is the interesting one: flat and ignored at ~3% for a decade, then 12.6%. That is a dormant idea from Japanese statistics finding its moment, and you can date the ignition in the archive.
The other constant is the community itself. Research Practice sits at 21–26% in every single era. These authors have always talked about themselves at exactly the same rate, through every paradigm shift. Make of that what you will.
3. The house method, confirmed
Eight independent readers, reading 540 essays between them without coordinating, converged on the same primary observation: the signature move is taking a philosophical question and finding a formalism that makes it tractable.
Garrabrant doing time and causality as set partitions because Pearl’s arrows won’t carry deterministic relationships. Eisenstat’s condensation — an information theory that optimises for how easily the encoding answers questions rather than for total code length. Kosoy building instrumental reward functions because an agent might care about paperclips it will never perceive. Hänni writing caring about people as a weighted adjacency matrix. Kelly betting derived as Nash bargaining among your possible future selves.
The second signature is the one the field should be proudest of: the self-attacking section. Posts that contain a heading naming the objection that breaks the post. Retractions in titles. Roughly one substantive essay in eight carries a public reversal, a correction credited by name to whoever supplied it, or a published null result — and the edit is left visible as a scar rather than silently merged.
That is genuinely rare. Almost nowhere else in professional discourse is the correction treated as data worth preserving.
4. Platform predicts length. Author predicts voice.
Medians: X 159–193 chars, LW comment 610–830, LW post 5,700–10,000.
But almost nobody modulates register across platforms. Jörn Stöhler posts only on X and writes pure LessWrong there — numbered reconstructions of an opponent’s model, raw odds notation in a tweet — though the data notes flag his X handle as the study’s one unconfirmed attribution. Wentworth posts only on LessWrong, zero X posts across all three samples, in a register so flat it does not change between Solomonoff induction, double-entry bookkeeping, fashion, and his own genome.
What X actually removes is blockquoting, footnotes, acknowledgements, and the ability to edit. Not personality.
And the strongest predictor of how much someone hedges turns out to be neither platform nor author: it is whether the claim is checkable. Technical claims get asserted. Community-strategy and forecasting claims get buried in qualifications. Which is, again, correct behaviour.
5. Things you should know before using this data
If you scrape this corpus — and people will — three things will wreck your statistics.
Up to a quarter of “documents” are not documents. Job ads. Meetup notices from 2014 with the address of a Del Taco. A publications index last updated in April 2000. A Netlify CMS config file, 14,500 characters of widget: string. One Lorem ipsum placeholder whose author could not remember the rest of the Lorem ipsum.
Fiction is 13.1% of the characters and almost all of it is one author’s Harry Potter fan fiction. Leave it in and the most distinctive term in twenty-three years of theoretical AI safety becomes Harry.
Institutional pages get attributed to people. 111 Schmidt Sciences press releases about climate modelling and ocean gyres sat under a researcher’s name in an intermediate build; an SFI release under Simon DeDeo’s; ~500 MIRI posts by other staff under Yudkowsky’s. All three are fixed in the final corpus — but they are exactly the failure mode to check for. A link to someone’s profile is not a link to their writing.
Also: LaTeX stripping systematically destroys exactly the documents you most want — the formal ones. Kosoy’s theorem posts survive extraction as “If , then .”
6. The outliers are where the community actually lives
The register that makes this community work is the same one that produces its strangest artefacts.
Wentworth investigating his own inability to feel companionate love by sequencing himself, finding a single-base-pair deletion in the oxytocin receptor ORF, and correctly noting which conclusions short-read sequencing cannot support. Confessional autobiography with a methods section.
Samin, after a Molotov cocktail was thrown at Sam Altman’s house, writing a flat deontological plea that ends by saying he wants absolutely everyone to survive, Altman included.
Yudkowsky’s eulogy for his grandfather, which opens by refuting a proverb about death on semantic grounds before grieving — and a separate post publicly asking to be falsified before he publishes a critique of OpenPhil. Adversarial collaboration as a genre.
Ngo publishing 2024 memos in 2026 with a preface saying he no longer thinks this kind of analysis is worth much.
And the one that should worry the field: Bushnaq mentioning in a shortform that his mathematics research is now largely done by AI agents working in parallel, and that the bottleneck is that they explain their ideas badly. That is the field’s subject matter arriving inside its daily practice, recorded the way you’d record the weather.
7. My actual opinions
The empirical turn is the best thing that has happened to this field, and it is not a betrayal of theory. It is the first development that made theory checkable. Causal scrubbing produces a number — and Chan published the one where the detailed hypothesis only recovered 72%. Local learning coefficients produce a number. Bushnaq published hand-coded weights that fall short of trained models by a factor of four and dared the community to beat him. The agent-foundations corpus of 2015–2020 is the most beautiful material in the archive and I’d not want it lost, but its decline is better read as the field acquiring the ability to be specifically wrong.
The real weakness is the missing middle. Wentworth’s The Plan – 2023 lists the median happy trajectory as: (1) sort out our confusions about agency and abstraction, (2) find a good alignment target and retarget the search, (3) …, (4) profit. The joke is deliberate. The joke is also the condition. Enormous care goes into building frameworks; very little into closing them. Open problems everywhere, endings that solicit rather than conclude, theorems published before their proofs. That is stable across all four eras — it is not a phase the field is passing through.
The flat analytical register does not generalise from arguments to people. Reduce the claim to a mechanism, deny the mechanism: superb against a bad argument, corrosive when the object is a colleague’s sexuality, a stranger’s body, or women in cost-benefit terms. There is a small but persistent vein of this — concentrated, by the social-corpus count, in one author, but too consistent to read as accident. It is not a rigour failure. It is rigour pointed somewhere it does not belong.
And the thing nobody is pricing in: concentration. One author is 19% of the long-form corpus by document count and 28% by characters. The top five are 58% of documents and 63% of characters. Whatever the community’s individual calibration norms are worth, a written record this dependent on a handful of voices is fragile in a way no amount of personal epistemic hygiene addresses.
LLM-to-read
Abstract
Synthesis of a two-corpus study of 49 theoretical-alignment (“ILIAD”) authors: 3,653 cleaned social posts (Aug 2025–Aug 2026) and 2,743 long-form documents (2003–2026, ~34.6M chars). The post reports that the community’s calibration practice is verbal rather than numeric; that long-form theme shares shifted sharply by era (deep learning, interpretability and evaluation up, agent foundations and rationality down, singular learning theory revived); that the signature intellectual move is formalising philosophical questions; that self-adversarial and visible-retraction norms are real and measurable; and that the corpus carries specific reuse hazards (non-documents, a large fiction share, repaired institutional misattribution). The author’s assessments: the post-2022 empirical turn made the theoretical programme checkable, the field under-invests in closing its frameworks, the flat analytical register is misapplied to persons in a small vein of writing, and the written record is heavily concentrated in a few authors.
Claims
Cross-corpus:
- C1. Theme distribution differs structurally between registers. Social is 30.8% conversational and 16.0% deep-learning; writings are 13.3% rationality, 10.1% agent foundations, 9.4% deep learning. The six themes the community is named for total 10.4% of social posts vs ~29% of writings. Social carries institutions and conflict; writings carry technical apparatus.
- C2. Calibration is verbal, not numeric, in both corpora. “Epistemic status” occurs 0–1 times per 244 social posts and in 2–9% of writings. Numeric credences about the world: ~2–4% of social posts, ~15–20% of writings. Verbal hedges outnumber numeric ones ~10:1. Post’s assessment: appropriate rather than hypocritical — numeric credences without models are noise.
- C3. Hedging is predicted by claim-checkability, not by platform or author. Technical claims are asserted; strategic and forecasting claims are heavily qualified.
- C4. Platform determines length; author determines voice. Author platform-choice is bimodal (several authors are 100% X, several 100% LessWrong). X removes blockquoting, footnotes, acknowledgements and editability without changing register.
- C5. The signature intellectual move is mathematising a philosophical problem. Independently reported as the primary observation by 8/8 writings analysts.
- C6. Self-adversarial norms are real and measurable. ~15% of substantive essays contain a section attacking the essay’s own proposal; ~1 in 8 contains a public reversal, named-credit correction, or published null result; ~3% of social posts contain a visible retraction left in place.
- C7. Analogy is the dominant proof strategy in both corpora, including within technical arguments.
Time-series (writings; era-banded share of pieces touching each theme):
| Theme | 2008–13 | 2014–18 | 2019–22 | 2023–26 | Direction |
|---|---|---|---|---|---|
| Deep Learning / LLMs | 9.4% | 13.4% | 32.4% | 43.4% | ↑ 4.6× |
| Mechanistic Interpretability | 1.6% | 1.7% | 7.1% | 20.0% | ↑ 12.5× |
| Evaluation, Oversight & Control | 1.9% | 4.2% | 5.6% | 17.2% | ↑ 9× |
| Singular Learning Theory | 6.1% | 2.9% | 3.3% | 12.6% | dormant → revival |
| AI Governance, Policy & Labs | 12.0% | 19.0% | 15.9% | 20.8% | ↑ |
| Agent Foundations & Decision Theory | 18.4% | 25.5% | 24.9% | 13.3% | peak 2014–18, then ↓48% |
| Rationality & Epistemics | 49.5% | 34.9% | 36.0% | 30.5% | monotonic ↓ |
| Research Practice & Community | 25.6% | 21.8% | 21.4% | 26.1% | flat |
- Formal correlates of the era shift: footnotes appear only post-2021; TL;DR/abstract blocks overwhelmingly post-2023; “Followup to:” pointers almost entirely pre-2012; acknowledgements, affiliation disclaimers and named-collaborator lists post-2022. No pre-2020 document reports an experiment the author ran on a model. Trajectory: forum-as-salon → forum-as-journal-plus-press-office.
- Social time-series (13 months, monthly): themes stable; two signals exceed noise — a Nov 2025 spike in Research Practice (20.9%), Rationality (21.2%) and Personal (11.0%), attributable in-text to the Inkhaven residency; and AI Governance rising through 2026 (5–7% → 10–13%).
Assessments (the post’s argued positions):
- A1. The 2022+ empirical turn made the theoretical programme checkable. Causal scrubbing, local learning coefficients and hand-coded-weight challenges produce beatable numbers. Read the agent-foundations decline as acquisition of falsifiability, not regression.
- A2. The dominant structural weakness is the missing middle: high framework-construction investment, low closure investment. Stable across all four eras; not a transitional phase.
- A3. The self-adversarial + visible-retraction norm is the corpus’s most transferable practice and is rare in professional discourse generally.
- A4. The flat analytical register generalises badly from propositions to persons. A small, consistent vein of writing — per the social-corpus reading, concentrated in one author — applies unchanged method to non-consenting third parties.
- A5. Concentration risk is unpriced: one author is 19.1% of documents and 28.1% of characters; the top five are 57.8% of documents and 62.6% of characters.
- A6. Highest-value under-examined datum: a working researcher reporting in passing that AI agents now perform the bulk of his mathematics, bottlenecked on their explanations. Subject matter has entered daily practice without being metabolised.
Data & provenance
- Social corpus:
social-media.json— 3,653 cleaned posts, 5.78M chars, 49 authors, 2025-08-29 → 2026-08-29; platforms X (2,048), LessWrong/Alignment Forum (1,851), EA Forum (18); X retrieved via a date-bounded commercial search API. - Writings corpus:
writings/*.json— 2,906 raw entries → 2,743 documents after dedup, ~34.6M chars, 2003–2026; scraped from links reachable from the study’s author index.
Method
Bottom-up lexical derivation of a 17-theme taxonomy, formalised as weighted phrase lexicons; length-normalised scoring with 100% coverage; monthly (social) and yearly (writings) time-series; a 20% theme-stratified sample (n=732 social, n=540 writings) read in full by independent analysts; 17 additional documents read directly for the opinion sections.
Reproduction
Pipeline (repo-relative): tools/corpus.py (clean, dedup, genre-tag, stable UIDs) → tools/themes.py (17-theme weighted-lexicon classifier) → tools/label.py (labels + time-series) → tools/sample.py (20% theme-stratified sample, agent bundles).
Artefacts:
| File | Contents |
|---|---|
themes-topics.txt | 17-theme taxonomy with definitions and both distributions |
social-media.md | Social corpus analysis (human + LLM sections) |
article-writings.md | Long-form corpus analysis (human + LLM sections) |
analysis/labeled_*.json | Per-document theme labels with stable UIDs |
analysis/timeseries.json | Monthly and yearly theme series |
analysis/sample_*.json, analysis/bundles/ | The 20% stratified sample and reader bundles |
tools/ | Full reproducible pipeline |
Caveats
Data-quality findings, critical for reuse (the post’s own):
- Between 1 in 12 and 1 in 4 writings “documents” is not authored prose: job ads, meetup notices, conference timetables, publications indices, seminar schedules, paper landing pages, one CMS
config.yml(~14.5K chars), one Lorem ipsum placeholder. Automated proxy floor 8.4%; reader hand-classification 12–25% per bundle. - Institutional-page misattribution — found and fixed upstream. 111 Schmidt Sciences press releases, 1 SFI press release, and ~500 MIRI-blog posts by other staff were present in an intermediate snapshot. The final scraper scope-locks profile-link crawls and drops feed entries bylined to others; the final corpus contains 0 institutional pages.
- Fiction is 13.1% of characters (166 docs, 4.52M chars), predominantly one author’s HPMOR chapters with a single import date. Segmented via a
genrefield. Unsegmented, it dominates term frequencies. - Date corruption in raw data: values of 0001, 2030, 2199; 182 undated (left empty, never guessed). Nulled outside 1980–2026.
- Co-authored pieces duplicate across author files by design (163 duplicate instances); dedup by URL/content hash required for corpus-level statistics.
- LaTeX stripping preferentially destroys formal documents, leaving semantically holed sentences; currency symbols stripped, producing nonsense figures.
- Theme labels are noisy at ~5–10%; all analysts independently flagged specific mislabels. Aggregate distributions are sound; individual labels are weak evidence.
- Social: X retrieved via a date-bounded commercial search API (deeper for quiet accounts than prolific ones); 5 accounts returned zero posts (unconfirmed, not verified silence); 1 X handle is unverified attribution and flagged in-file.
Editorial:
- The title’s “6,500 documents” is a loose round: the cleaned corpora total 6,396 documents (2,743 + 3,653); raw entries total 6,559 (2,906 + 3,653). Kept as the author wrote it.
- Characterisations of named researchers rest on their public posts as sampled here; one X-handle attribution is unconfirmed (caveat 8 above).
Provenance: edited September 2026.