Twenty-three years of a field talking itself into existence
One of three essays from a single corpus study of 49 theoretical-alignment (“ILIAD”) authors: this one reads the long-form archive (2,743 pieces, 2003–2026); its companions read the social corpus (3,653 posts, Aug 2025–Aug 2026) and the cross-corpus synthesis.
The long-form archive is a different animal from the social one. Where the posts are a room full of people interrupting each other, the writings are a library — 2,743 pieces, 34.6 million characters, reaching back to 2003, when the oldest document here was written and the phrase “AI alignment” did not yet mean anything to anyone.
Read them in order and you watch a discipline being built out of nothing by people who mostly did not have jobs doing it.
The arc is visible in the numbers and it is unambiguous. In 2008–2013, half of everything published (49.5%) is about rationality and epistemics — how to think, how to notice you are wrong, how to hold a belief loosely. Deep learning is 9.4%. Mechanistic interpretability is 1.6%, which is to say it does not exist. By 2023–2026 that has inverted: rationality has fallen to 30.5%, deep learning has climbed to 43.4%, interpretability has gone up twelvefold to 20.0%, and evaluation and oversight — safety cases, red-teaming, control — has gone up ninefold to 17.2%.
And in the middle of that arc sits the story the field does not much like to tell about itself. Agent foundations — decision theory, logical induction, embedded agency, Cartesian frames, the whole magnificent MIRI edifice — climbs from 18.4% to a peak of 25.5% in 2014–2018, holds at 24.9%, and then falls to 13.3%. The most beautiful mathematics this community produced is now the smallest it has ever been as a share of what it writes.
Singular learning theory does the opposite. It sits at 2.9% and 3.3% through the middle years, ignored, and then jumps to 12.6% in 2023–2026. A dormant idea from Japanese statistics, waiting for someone to notice that neural networks are singular models. Daniel Murfet and Jesse Hoogland noticed, and you can watch the ignition happen in the archive.
What the writing is actually like
There is one intellectual move that recurs so consistently across authors, decades and venues that it is fair to call it the house method: take a philosophical question and find a formalism that makes it tractable.
Scott Garrabrant, wanting to understand time and causality without Pearl’s arrows, builds finite factored sets out of set partitions and starts a talk by admitting that if he were better he would be giving a category theory talk instead, but he trained as a combinatorialist. Sam Eisenstat asks what a concept is, and answers with “condensation” — an information theory that optimises not for total code length but for how easily the encoding answers questions about the data. Compression creates density; condensation forms discrete droplets. Vanessa Kosoy notices that an agent might care about things it cannot perceive — a donation to a malaria charity, a paperclip on the far side of the universe — and builds “instrumental reward functions” to make regret analysis honest about it. Kaarel Hänni takes caring about other people and writes it as a weighted adjacency matrix.
Each of these is a philosophical problem that a philosopher would have written an essay about, converted into an object with theorems attached. Sometimes it works spectacularly. Sometimes, as several of these authors say themselves, it produces a beautiful structure that nothing has yet been built on.
The second signature is stranger and more admirable: the author attacks his own proposal, in-line, under a heading. Posts contain sections titled with the objection that breaks the thing the post just proposed. But What If There’s An External Adversary? Is This Just The Markov Blanket? One piece has its own retraction in its title. Another opens by disowning the post before it. Roughly one substantive essay in eight contains a public reversal, a correction credited by name to the person who supplied it, or a published null result. The credibility currency in this community is not being right; it is demonstrated adversarial pressure on your own idea.
And then the third thing, which no one designed: this is a corpus where a formal theorem and a confession live in the same venue, under the same byline, in the same voice. A researcher publishes a proof about optimal predictors and, months later, a public letter about the end of a ten-year relationship. Another posts a career self-audit admitting the year felt slow and unproductive and naming his supervisor. Another synthesises a COVID vaccine in his kitchen from an open-source design and posts the costs. The personal and the professional are not separated because it never occurred to anyone to separate them.
What the archive is not
Here is the finding that matters most for anyone who intends to use this data, and it emerged independently from all eight readers: somewhere between one document in twelve and one in four is not a piece of writing at all.
Job advertisements. Meetup notices from 2014 with the address of a Del Taco. Conference timetables. A university publications index last updated in April 2000. A seminar schedule. A paper landing page. A Netlify CMS configuration file — 14,500 characters of widget: string declarations. And, my favourite, a placeholder post containing Lorem ipsum and an admission that its author could not remember the rest of the Lorem ipsum.
These are scraper artefacts, and they sit in the archive wearing the same metadata as a theorem. Any statistic computed over the raw files without filtering them is wrong. An automated proxy (very short documents plus boilerplate markers) puts the floor at 8.4%; hand classification by readers puts it higher, at one in eight to one in four depending on how strictly you define a non-document.
Two further contaminations of the same kind were caught and fixed during scraping, so they are absent from the final corpus: 111 Schmidt Sciences press releases about climate modelling and ocean gyres, harvested from an institutional newsroom and attributed to a researcher who did not write them; and roughly 500 MIRI-blog posts by other staff sitting in one author’s file. Both were present in an intermediate snapshot. The lesson generalises: a link to someone’s profile on an institution’s site is not a link to their writing.
The same applies to genre. 4.52 million characters of this corpus — 13.1% — is fiction, and almost all of it is one author’s Harry Potter fan fiction, scraped as chapters with a single import date. It is genuinely his online writing. It is also 660,000 words of a novel sitting inside a corpus about alignment research, and if you let it into your term frequencies, the most distinctive word in twenty-three years of theoretical AI safety becomes Harry.
My opinion
I read seventeen of these pieces end to end, chosen to span every theme and twenty different authors, and my honest reaction is admiration braided with a specific frustration.
The admiration is for the seriousness. Kosoy’s regret bound, Chan’s causal scrubbing results reporting that the coarse hypothesis recovered 88–93% of loss but the detailed one only 72% — and publishing the 72% — Bushnaq’s hand-coded MLP weights that memorise facts at a scaling prefactor still four times worse than trained models, published as an open challenge to beat him. That is real work with falsifiable content, and the willingness to publish the number that makes you look worse is not common in most fields.
The frustration is the missing middle. John Wentworth’s “The Plan – 2023 Version” lays out the median happy trajectory as: (1) sort out our fundamental confusions about agency and abstraction, (2) find a good alignment target in the AI’s internal concepts and retarget the search, (3) …, (4) profit. The joke is deliberate and self-aware — he knows exactly what he is doing — but the joke is also the field’s actual condition. There is an enormous amount of care spent constructing frameworks and remarkably little spent closing them. Open problems everywhere. Endings that solicit rather than conclude. A 2018 theorem post that admits the proof is not written yet and that every attempt to extend it has failed, published anyway, which is honest and admirable and also, twenty-three years in, a pattern.
My view is that the empirical turn visible after 2022 is not a betrayal of the theoretical project, as some of these authors clearly fear. It is the first thing that has made the theoretical project checkable. Causal scrubbing gives a number. The local learning coefficient gives a number. Bushnaq’s challenge gives a number you can beat. The agent-foundations work of 2015–2020 is the most intellectually beautiful material in this archive and I would not want it lost — but its decline from 25.5% to 13.3% is not obviously the field getting dumber. It might be the field acquiring the ability to be wrong about something specific.
The thing I would actually worry about is concentration. One author is a fifth of this corpus by documents and more than a quarter by characters. The five most prolific are well over half of both. A field whose written record is this dependent on a handful of voices has a fragility that no amount of individual calibration fixes.
LLM-to-read
Abstract
Analysis of the long-form corpus of a study of 49 theoretical-alignment (“ILIAD”) authors: 2,743 documents (~34.6M chars, 2003–2026) after cleaning and deduplication. Era-banded theme shares document an empirical turn (deep learning 9.4%→43.4%, mechanistic interpretability 1.6%→20.0%, evaluation/oversight 1.9%→17.2%), the peak and decline of agent foundations (25.5%→13.3%), a singular-learning-theory revival (~3%→12.6%) and a monotonic decline of rationality writing (49.5%→30.5%). Qualitative reading of a 20% stratified sample identifies the signature intellectual move (mathematising philosophical problems), self-adversarial sectioning, visible retraction, toy-model evidence and analogy-driven argument. The post also documents reuse hazards: non-document contamination, a large single-author fiction share, upstream-fixed institutional misattribution and LaTeX-stripping damage. Assessments: the empirical turn made the theoretical programme checkable; the field’s dominant weakness is a “missing middle” of unclosed frameworks; the written record is heavily concentrated in a few authors.
Claims
Primary-theme distribution (share of 2,743 pieces; same 17-theme taxonomy as the social corpus, coverage 100%):
| Theme | Share |
|---|---|
| Rationality, Epistemics & Forecasting | 13.3% |
| Agent Foundations & Decision Theory | 10.1% |
| Deep Learning, LLMs & Capabilities | 9.4% |
| Research Practice, Career & Community | 8.9% |
| Personal, Culture & Miscellany | 8.2% |
| Fiction & Narrative | 7.8% |
| Pure Mathematics | 6.9% |
| Biology, Complexity & Multi-Agent Systems | 6.8% |
| Philosophy, Ethics & Consciousness | 4.5% |
| AI Governance, Policy & Labs | 4.4% |
| Singular Learning Theory & Loss Landscape | 4.2% |
| Alignment Threat Models & Failure Modes | 4.1% |
| Mechanistic Interpretability | 3.5% |
| Natural Abstraction & Latent Structure | 2.8% |
| Evaluation, Oversight & Control | 2.4% |
| Computational Mechanics & Information Theory | 1.5% |
| Conversational & Reactive | 1.4% |
- Distinctive vocabulary vs the social corpus (log-odds):
mesa-optimizer,cartesian,morphism,subagent,factored,goodhart,polytope,counterfactuals,orthogonality,partition,proposition. The writings carry the technical apparatus; the social corpus carries the institutions and the gossip.
Time-series (share of pieces touching each theme, era-banded):
| Theme | 2008–13 (n=309) | 2014–18 (n=648) | 2019–22 (n=747) | 2023–26 (n=855) |
|---|---|---|---|---|
| Rationality, Epistemics & Forecasting | 49.5% | 34.9% | 36.0% | 30.5% |
| Agent Foundations & Decision Theory | 18.4% | 25.5% | 24.9% | 13.3% |
| Deep Learning, LLMs & Capabilities | 9.4% | 13.4% | 32.4% | 43.4% |
| Mechanistic Interpretability | 1.6% | 1.7% | 7.1% | 20.0% |
| Evaluation, Oversight & Control | 1.9% | 4.2% | 5.6% | 17.2% |
| Singular Learning Theory & Loss Landscape | 6.1% | 2.9% | 3.3% | 12.6% |
| AI Governance, Policy & Labs | 12.0% | 19.0% | 15.9% | 20.8% |
| Pure Mathematics | 12.3% | 25.9% | 15.8% | 13.1% |
| Research Practice, Career & Community | 25.6% | 21.8% | 21.4% | 26.1% |
Four time-series findings:
- The empirical turn. Deep learning 9.4%→43.4% (4.6×); mech interp 1.6%→20.0% (12.5×); evals/oversight 1.9%→17.2% (9×). No document before ~2020 reports an experiment the author ran on a model; by 2023–26 this is a standard genre.
- Agent foundations peaked and declined: 25.5% (2014–18) → 13.3% (2023–26), a 48% fall from peak.
- SLT revival: dormant at ~3% through 2014–2022, then 12.6%. Traceable to specific authors (Murfet, Hoogland) and an institution (Timaeus, later merged into Resolution).
- Rationality declines monotonically 49.5%→30.5% as the field professionalises. Research Practice is flat (~21–26%) across all four eras — the community has always talked about itself at a constant rate.
- Formal correlates of the era shift: footnote markers appear only from ~2021–22 onward; bolded TL;DR/abstract blocks are overwhelmingly post-2023; “Followup to:” sequence pointers are almost entirely pre-2012; acknowledgements paragraphs, institutional affiliation disclaimers and named collaborator lists are post-2022.
Common features (full reading of a 20% theme-stratified sample, n=540, 8 independent readers):
- Signature intellectual move: mathematise a philosophical or social problem. Consistently the most-reported feature across all eight readers. Examples: causality without Pearl (finite factored sets); concept-formation as an encoding problem (condensation); “purpose” in terms of maths and physics; Kelly betting as Nash bargaining among possible future selves; interpersonal caring as a weighted adjacency matrix.
- Self-supplied counterexample, under its own heading. Posts contain named sections that attack the proposal the post just made. ~15% of substantive essays.
- Public retraction and in-place revision. ~1 in 8 substantive essays contains a visible reversal, a correction credited to a named corrector, or a published null result. Edits are left as scars, not silently merged.
- Toy models are the standard unit of evidence, deliberately banal and fully solvable: two coins, a marble in a double bowl, a ripple-carry adder, bleggs and rubes, a roadtrip and a screwdriver.
- Cross-domain example lists establish that a concept is “natural” by exhibiting it in physics, engineering, biology and everyday speech before any formalism appears.
- Diagrams-in-words. Figures did not survive extraction, exposing how heavily these authors narrate images. ~10–15% of docs contain dangling deixis referring to absent figures.
- Citation is social, not bibliographic. Inline hyperlinks (now flattened) and first-name references to community members. Formal author-year citation appears only in the academic-adjacent clusters (Timaeus/SLT, evals papers, Sandberg).
- Openings: meta-preamble about the document’s own status/provenance (~25%); explicit dependency pointer to a prior post; or a concrete particular before any abstraction.
- Explicit “Epistemic status:” headers are rare — 1 to 6 per 68-document bundle (~2–9%), far below the convention’s reputation. Function is served instead by italic preambles and dense inline hedging.
- Closings are open: explicit “Open Problem” sections, requests for reader data, or a punt.
Patterns in the differences:
- By era: 2008–2013 is conversational, aphoristic, second-person, formalism-free, and administratively small-scale (meetup notices, site announcements). 2015–2020 is agent-foundations-as-sequence: Definition/Proposition/Proof with minimal motivation, written for dozens of readers. 2022–2026 is institutional: TL;DRs, acknowledgements, affiliation disclaimers, benchmarks, hiring pages. Trajectory: forum-as-salon → forum-as-journal-plus-press-office.
- By venue: LessWrong (~72–85% of docs) assumes shared canon and uses platform affordances (spoiler tags, hashed footnotes, sequence links). Personal blogs are lower-stakes, more disciplinary, more outward-facing. Institutional sites (timaeus.co, intelligence.org) contain nearly no argument — job ads, project listings, press copy.
- By author (representative contrasts): Kosoy is the formal extreme, offering no motivation or worked examples; Demski is the most procedural and most self-revising; Garrabrant is the most compressed (mean ~2,570 chars vs corpus ~11,200) and most willing to publish incomplete work labelled incomplete; Wentworth has the widest genre range in the corpus and the flattest register across it; Ngo is the analytic philosopher and the corpus’s most accomplished fiction writer; Sandberg belongs to a visibly different tradition (Oxford futures studies, recreational computational mathematics, no LW jargon, no credences).
- By genre, which predicts prose better than theme does: research report, theorem paper, tutorial, polemic, dialogue/transcript, fiction, and administrative notice barely share conventions.
Outliers:
- Non-documents (~1 in 6): job ads, meetup notices, conference timetables, a publications index last updated April 2000, a seminar schedule, paper landing pages, a Netlify/Decap CMS
config.yml(~14,500 chars), and one Lorem ipsum placeholder. Filter before computing any statistic. - Dialogue/transcript as a primary research artefact: timestamped chat logs between named researchers published unedited, preserving live disagreement and mid-argument mind-changes. ~11% of one bundle; strongly associated with one author.
- Confessional material in a research venue: a public letter about a relationship’s end; a PhD self-audit naming a supervisor; a self-sequencing investigation into the author’s own capacity for attachment; a kitchen-synthesised COVID vaccine with costs.
- Fiction carrying doctrine: a first-person monologue in the voice of an aligned superintelligence; Aesop’s fable rewritten ~12 times as successive EA failure modes; a 2003 story about an NPC who realises she is one and asks to be made real.
- Intra-field conflict published on the record: a post arguing the AI safety community is structurally power-seeking, written from inside a frontier lab with a disclaimer that it was not reviewed; a direct claim that a load-bearing statement repeated across SLT papers is wrong; a public reversal on a research agenda credited to a single conversation, with the author noting he may have updated partly because he respects the interlocutor and calling that “particularly embarrassing.”
- Extraction damage: LaTeX stripped from formal pieces leaves semantically holed sentences (“If , then .”); currency symbols stripped, producing nonsense figures; one arXiv conversion degenerates into raw TikZ internals. Formal documents are systematically the most damaged.
Assessments (the post’s argued positions):
- The empirical turn (2022→) is the first development that made the theoretical programme checkable. Causal scrubbing, local learning coefficients, and hand-coded-weight challenges all produce numbers that can be beaten. The decline of agent foundations from 25.5% to 13.3% is better read as the field acquiring falsifiability than as intellectual regression.
- The corpus’s dominant weakness is the missing middle: high investment in framework construction, low investment in closure. Open problems, soliciting endings, and theorems published before their proofs are a stable pattern across all four eras, not a phase.
- The self-adversarial norm (attack your own proposal under a heading; leave the retraction visible) is the most transferable practice in the archive.
- Concentration risk: one author is 19.1% of documents and 28.1% of characters; the top five are 57.8% of documents and 62.6% of characters. The written record of this field rests on very few people.
Data & provenance
- Source:
writings/<author>.json, 49 files, 2,906 raw entries (final scraper output); reachable from the links inauthors.json. Not a complete bibliography — it is what those links reached. - After cleaning and deduplication: 2,743 documents, ~34.6M chars, 2003–2026.
- Deduplication: co-authored pieces appear in multiple author files by design; deduplicated by URL/content hash for corpus-level statistics (163 duplicate instances removed), with co-authorship recorded.
- Segmented out of the analysable set: 166 fiction documents (4.52M chars, 13.1% of characters, predominantly HPMOR chapters).
- Fixed upstream during scraping: 111 Schmidt Sciences press releases and one SFI press release (institutional newsroom pages misattributed to researchers), plus ~500 MIRI-blog posts by other staff in one author’s file. These were present in an intermediate snapshot and are absent from the final corpus; institutional-page count is now 0.
- Cleaning: URLs, HTML, CDN/CSS artefacts, LaTeX commands, markdown image refs stripped. Dates outside 1980–2026 nulled (raw data contained dates of 0001, 2030, 2199).
Method
Same 17-theme taxonomy as the social corpus, derived bottom-up from term/bigram frequency and formalised as weighted phrase lexicons; every document scored length-normalised, coverage 100%. Era-banded time-series over 2008–2026. A 20% theme-stratified sample (n=540) was read in full by 8 independent readers; the author additionally read 17 pieces end to end, spanning every theme, for the opinion section.
Reproduction
- Taxonomy:
themes-topics.txt. Corpus files:writings/<author>.json; author index:authors.json. - Pipeline: the study’s
tools/directory (cleaning, dedup, labelling, sampling — documented in the synthesis essay).
Caveats
The post’s own:
- Theme labels are automatic and noisy at roughly the 5–10% level; all eight readers independently flagged specific mislabels. The stratified sample was drawn from a slightly earlier snapshot (2,856 docs) than the final corpus (2,743); qualitative findings are unaffected, and all quantitative figures above are recomputed on the final data. Treat any single label as weak evidence, and the aggregate distribution as sound.
- 8 of 49 authors have empty archives (Hsia, Little, Abrams, Aoyagi, Griffin, Dell, Rosas, Shankar) — all genuine absences of long-form writing outside excluded platforms, not blocks or broken feeds.
- The corpus is not a complete bibliography: it is what was reachable from the links in
authors.json. - Character-count statistics are distorted by fiction unless segmented; document-count statistics are distorted by non-documents unless filtered.
Editorial:
- Era bands cover 2008–2026; earlier documents (the corpus reaches back to 2003) and undated ones fall outside the banded columns.
- The human section’s CMS-file figure originally read “14,000 characters”; harmonised to the ~14,500 given by this post’s own reference material and the companion essays.
- Confessional documents are referenced without names here, but several may be identifiable to specialist readers. Consent questions around them remain open.
Provenance: edited September 2026.