Foomax

Twenty-three years of a field talking itself into existence

September 2026 · ILIAD

One of three essays from a single corpus study of 49 theoretical-alignment (“ILIAD”) authors: this one reads the long-form archive (2,743 pieces, 2003–2026); its companions read the social corpus (3,653 posts, Aug 2025–Aug 2026) and the cross-corpus synthesis.

The long-form archive is a different animal from the social one. Where the posts are a room full of people interrupting each other, the writings are a library — 2,743 pieces, 34.6 million characters, reaching back to 2003, when the oldest document here was written and the phrase “AI alignment” did not yet mean anything to anyone.

Read them in order and you watch a discipline being built out of nothing by people who mostly did not have jobs doing it.

The arc is visible in the numbers and it is unambiguous. In 2008–2013, half of everything published (49.5%) is about rationality and epistemics — how to think, how to notice you are wrong, how to hold a belief loosely. Deep learning is 9.4%. Mechanistic interpretability is 1.6%, which is to say it does not exist. By 2023–2026 that has inverted: rationality has fallen to 30.5%, deep learning has climbed to 43.4%, interpretability has gone up twelvefold to 20.0%, and evaluation and oversight — safety cases, red-teaming, control — has gone up ninefold to 17.2%.

And in the middle of that arc sits the story the field does not much like to tell about itself. Agent foundations — decision theory, logical induction, embedded agency, Cartesian frames, the whole magnificent MIRI edifice — climbs from 18.4% to a peak of 25.5% in 2014–2018, holds at 24.9%, and then falls to 13.3%. The most beautiful mathematics this community produced is now the smallest it has ever been as a share of what it writes.

Singular learning theory does the opposite. It sits at 2.9% and 3.3% through the middle years, ignored, and then jumps to 12.6% in 2023–2026. A dormant idea from Japanese statistics, waiting for someone to notice that neural networks are singular models. Daniel Murfet and Jesse Hoogland noticed, and you can watch the ignition happen in the archive.

What the writing is actually like

There is one intellectual move that recurs so consistently across authors, decades and venues that it is fair to call it the house method: take a philosophical question and find a formalism that makes it tractable.

Scott Garrabrant, wanting to understand time and causality without Pearl’s arrows, builds finite factored sets out of set partitions and starts a talk by admitting that if he were better he would be giving a category theory talk instead, but he trained as a combinatorialist. Sam Eisenstat asks what a concept is, and answers with “condensation” — an information theory that optimises not for total code length but for how easily the encoding answers questions about the data. Compression creates density; condensation forms discrete droplets. Vanessa Kosoy notices that an agent might care about things it cannot perceive — a donation to a malaria charity, a paperclip on the far side of the universe — and builds “instrumental reward functions” to make regret analysis honest about it. Kaarel Hänni takes caring about other people and writes it as a weighted adjacency matrix.

Each of these is a philosophical problem that a philosopher would have written an essay about, converted into an object with theorems attached. Sometimes it works spectacularly. Sometimes, as several of these authors say themselves, it produces a beautiful structure that nothing has yet been built on.

The second signature is stranger and more admirable: the author attacks his own proposal, in-line, under a heading. Posts contain sections titled with the objection that breaks the thing the post just proposed. But What If There’s An External Adversary? Is This Just The Markov Blanket? One piece has its own retraction in its title. Another opens by disowning the post before it. Roughly one substantive essay in eight contains a public reversal, a correction credited by name to the person who supplied it, or a published null result. The credibility currency in this community is not being right; it is demonstrated adversarial pressure on your own idea.

And then the third thing, which no one designed: this is a corpus where a formal theorem and a confession live in the same venue, under the same byline, in the same voice. A researcher publishes a proof about optimal predictors and, months later, a public letter about the end of a ten-year relationship. Another posts a career self-audit admitting the year felt slow and unproductive and naming his supervisor. Another synthesises a COVID vaccine in his kitchen from an open-source design and posts the costs. The personal and the professional are not separated because it never occurred to anyone to separate them.

What the archive is not

Here is the finding that matters most for anyone who intends to use this data, and it emerged independently from all eight readers: somewhere between one document in twelve and one in four is not a piece of writing at all.

Job advertisements. Meetup notices from 2014 with the address of a Del Taco. Conference timetables. A university publications index last updated in April 2000. A seminar schedule. A paper landing page. A Netlify CMS configuration file — 14,500 characters of widget: string declarations. And, my favourite, a placeholder post containing Lorem ipsum and an admission that its author could not remember the rest of the Lorem ipsum.

These are scraper artefacts, and they sit in the archive wearing the same metadata as a theorem. Any statistic computed over the raw files without filtering them is wrong. An automated proxy (very short documents plus boilerplate markers) puts the floor at 8.4%; hand classification by readers puts it higher, at one in eight to one in four depending on how strictly you define a non-document.

Two further contaminations of the same kind were caught and fixed during scraping, so they are absent from the final corpus: 111 Schmidt Sciences press releases about climate modelling and ocean gyres, harvested from an institutional newsroom and attributed to a researcher who did not write them; and roughly 500 MIRI-blog posts by other staff sitting in one author’s file. Both were present in an intermediate snapshot. The lesson generalises: a link to someone’s profile on an institution’s site is not a link to their writing.

The same applies to genre. 4.52 million characters of this corpus — 13.1% — is fiction, and almost all of it is one author’s Harry Potter fan fiction, scraped as chapters with a single import date. It is genuinely his online writing. It is also 660,000 words of a novel sitting inside a corpus about alignment research, and if you let it into your term frequencies, the most distinctive word in twenty-three years of theoretical AI safety becomes Harry.

My opinion

I read seventeen of these pieces end to end, chosen to span every theme and twenty different authors, and my honest reaction is admiration braided with a specific frustration.

The admiration is for the seriousness. Kosoy’s regret bound, Chan’s causal scrubbing results reporting that the coarse hypothesis recovered 88–93% of loss but the detailed one only 72% — and publishing the 72% — Bushnaq’s hand-coded MLP weights that memorise facts at a scaling prefactor still four times worse than trained models, published as an open challenge to beat him. That is real work with falsifiable content, and the willingness to publish the number that makes you look worse is not common in most fields.

The frustration is the missing middle. John Wentworth’s “The Plan – 2023 Version” lays out the median happy trajectory as: (1) sort out our fundamental confusions about agency and abstraction, (2) find a good alignment target in the AI’s internal concepts and retarget the search, (3) , (4) profit. The joke is deliberate and self-aware — he knows exactly what he is doing — but the joke is also the field’s actual condition. There is an enormous amount of care spent constructing frameworks and remarkably little spent closing them. Open problems everywhere. Endings that solicit rather than conclude. A 2018 theorem post that admits the proof is not written yet and that every attempt to extend it has failed, published anyway, which is honest and admirable and also, twenty-three years in, a pattern.

My view is that the empirical turn visible after 2022 is not a betrayal of the theoretical project, as some of these authors clearly fear. It is the first thing that has made the theoretical project checkable. Causal scrubbing gives a number. The local learning coefficient gives a number. Bushnaq’s challenge gives a number you can beat. The agent-foundations work of 2015–2020 is the most intellectually beautiful material in this archive and I would not want it lost — but its decline from 25.5% to 13.3% is not obviously the field getting dumber. It might be the field acquiring the ability to be wrong about something specific.

The thing I would actually worry about is concentration. One author is a fifth of this corpus by documents and more than a quarter by characters. The five most prolific are well over half of both. A field whose written record is this dependent on a handful of voices has a fragility that no amount of individual calibration fixes.

LLM-to-read

Abstract

Analysis of the long-form corpus of a study of 49 theoretical-alignment (“ILIAD”) authors: 2,743 documents (~34.6M chars, 2003–2026) after cleaning and deduplication. Era-banded theme shares document an empirical turn (deep learning 9.4%→43.4%, mechanistic interpretability 1.6%→20.0%, evaluation/oversight 1.9%→17.2%), the peak and decline of agent foundations (25.5%→13.3%), a singular-learning-theory revival (~3%→12.6%) and a monotonic decline of rationality writing (49.5%→30.5%). Qualitative reading of a 20% stratified sample identifies the signature intellectual move (mathematising philosophical problems), self-adversarial sectioning, visible retraction, toy-model evidence and analogy-driven argument. The post also documents reuse hazards: non-document contamination, a large single-author fiction share, upstream-fixed institutional misattribution and LaTeX-stripping damage. Assessments: the empirical turn made the theoretical programme checkable; the field’s dominant weakness is a “missing middle” of unclosed frameworks; the written record is heavily concentrated in a few authors.

Claims

Primary-theme distribution (share of 2,743 pieces; same 17-theme taxonomy as the social corpus, coverage 100%):

ThemeShare
Rationality, Epistemics & Forecasting13.3%
Agent Foundations & Decision Theory10.1%
Deep Learning, LLMs & Capabilities9.4%
Research Practice, Career & Community8.9%
Personal, Culture & Miscellany8.2%
Fiction & Narrative7.8%
Pure Mathematics6.9%
Biology, Complexity & Multi-Agent Systems6.8%
Philosophy, Ethics & Consciousness4.5%
AI Governance, Policy & Labs4.4%
Singular Learning Theory & Loss Landscape4.2%
Alignment Threat Models & Failure Modes4.1%
Mechanistic Interpretability3.5%
Natural Abstraction & Latent Structure2.8%
Evaluation, Oversight & Control2.4%
Computational Mechanics & Information Theory1.5%
Conversational & Reactive1.4%

Time-series (share of pieces touching each theme, era-banded):

Theme2008–13 (n=309)2014–18 (n=648)2019–22 (n=747)2023–26 (n=855)
Rationality, Epistemics & Forecasting49.5%34.9%36.0%30.5%
Agent Foundations & Decision Theory18.4%25.5%24.9%13.3%
Deep Learning, LLMs & Capabilities9.4%13.4%32.4%43.4%
Mechanistic Interpretability1.6%1.7%7.1%20.0%
Evaluation, Oversight & Control1.9%4.2%5.6%17.2%
Singular Learning Theory & Loss Landscape6.1%2.9%3.3%12.6%
AI Governance, Policy & Labs12.0%19.0%15.9%20.8%
Pure Mathematics12.3%25.9%15.8%13.1%
Research Practice, Career & Community25.6%21.8%21.4%26.1%

Four time-series findings:

  1. The empirical turn. Deep learning 9.4%→43.4% (4.6×); mech interp 1.6%→20.0% (12.5×); evals/oversight 1.9%→17.2% (9×). No document before ~2020 reports an experiment the author ran on a model; by 2023–26 this is a standard genre.
  2. Agent foundations peaked and declined: 25.5% (2014–18) → 13.3% (2023–26), a 48% fall from peak.
  3. SLT revival: dormant at ~3% through 2014–2022, then 12.6%. Traceable to specific authors (Murfet, Hoogland) and an institution (Timaeus, later merged into Resolution).
  4. Rationality declines monotonically 49.5%→30.5% as the field professionalises. Research Practice is flat (~21–26%) across all four eras — the community has always talked about itself at a constant rate.

Common features (full reading of a 20% theme-stratified sample, n=540, 8 independent readers):

  1. Signature intellectual move: mathematise a philosophical or social problem. Consistently the most-reported feature across all eight readers. Examples: causality without Pearl (finite factored sets); concept-formation as an encoding problem (condensation); “purpose” in terms of maths and physics; Kelly betting as Nash bargaining among possible future selves; interpersonal caring as a weighted adjacency matrix.
  2. Self-supplied counterexample, under its own heading. Posts contain named sections that attack the proposal the post just made. ~15% of substantive essays.
  3. Public retraction and in-place revision. ~1 in 8 substantive essays contains a visible reversal, a correction credited to a named corrector, or a published null result. Edits are left as scars, not silently merged.
  4. Toy models are the standard unit of evidence, deliberately banal and fully solvable: two coins, a marble in a double bowl, a ripple-carry adder, bleggs and rubes, a roadtrip and a screwdriver.
  5. Cross-domain example lists establish that a concept is “natural” by exhibiting it in physics, engineering, biology and everyday speech before any formalism appears.
  6. Diagrams-in-words. Figures did not survive extraction, exposing how heavily these authors narrate images. ~10–15% of docs contain dangling deixis referring to absent figures.
  7. Citation is social, not bibliographic. Inline hyperlinks (now flattened) and first-name references to community members. Formal author-year citation appears only in the academic-adjacent clusters (Timaeus/SLT, evals papers, Sandberg).
  8. Openings: meta-preamble about the document’s own status/provenance (~25%); explicit dependency pointer to a prior post; or a concrete particular before any abstraction.
  9. Explicit “Epistemic status:” headers are rare — 1 to 6 per 68-document bundle (~2–9%), far below the convention’s reputation. Function is served instead by italic preambles and dense inline hedging.
  10. Closings are open: explicit “Open Problem” sections, requests for reader data, or a punt.

Patterns in the differences:

Outliers:

Assessments (the post’s argued positions):

  1. The empirical turn (2022→) is the first development that made the theoretical programme checkable. Causal scrubbing, local learning coefficients, and hand-coded-weight challenges all produce numbers that can be beaten. The decline of agent foundations from 25.5% to 13.3% is better read as the field acquiring falsifiability than as intellectual regression.
  2. The corpus’s dominant weakness is the missing middle: high investment in framework construction, low investment in closure. Open problems, soliciting endings, and theorems published before their proofs are a stable pattern across all four eras, not a phase.
  3. The self-adversarial norm (attack your own proposal under a heading; leave the retraction visible) is the most transferable practice in the archive.
  4. Concentration risk: one author is 19.1% of documents and 28.1% of characters; the top five are 57.8% of documents and 62.6% of characters. The written record of this field rests on very few people.

Data & provenance

Method

Same 17-theme taxonomy as the social corpus, derived bottom-up from term/bigram frequency and formalised as weighted phrase lexicons; every document scored length-normalised, coverage 100%. Era-banded time-series over 2008–2026. A 20% theme-stratified sample (n=540) was read in full by 8 independent readers; the author additionally read 17 pieces end to end, spanning every theme, for the opinion section.

Reproduction

Caveats

The post’s own:

Editorial:

Provenance: edited September 2026.