Foomax

The Headless Chooks at the Hinge

September 2026

The AI-safety movement saw a real danger. Then it mistook seeing first for the right to steer.

There is a true thing at the heart of AI safety.

A machine more capable than its makers could do terrible damage. It need not hate us. It needs reach, a bad objective and people foolish enough to give it authority. We already build systems we do not understand. We already deploy them because the quarter is ending.

That part is true.

Then there is the drug.

If you believe you noticed the decisive danger before the rest of humanity, you become important. If this century is the hinge of history, your seminar is no longer a seminar. It is the war room of the species. Your research agenda is not one proposal among many. It is the plan. Your career is not a career. It is the thin line between the light cone and the grave.

That is strong medicine for clever, frightened people.

The capital-S AI-safety movement was built from the true thing and the drug. It has never learned to separate them.

By the movement I do not mean every engineer working on reliability, cybersecurity, privacy, bias, accident prevention or product assurance. Those are fields. They have tools, customers and corpses they can count.

I mean the family descended from rationalism, singularitarianism, effective altruism, longtermism, superintelligence and existential risk. It grew around Eliezer Yudkowsky and MIRI, Nick Bostrom and Oxford’s Future of Humanity Institute, LessWrong, the Centre for Effective Altruism and a small donor network that believed artificial intelligence might end the human story.

They had no superintelligence to study. They had arguments.

So they imagined an agent with a clean utility function. They gave it an innocent goal. They made it strong. The innocent goal ate the world. They named the moves: orthogonality, instrumental convergence, corrigibility, value lock-in, the vulnerable world, the unilateralist’s curse.

Some of those were excellent ideas. They made an invisible class of danger visible. They also gave a field founded before its object existed the habits of theology. A parable stood in for an experiment. A possibility became a scenario. The scenario acquired a probability. The probability entered an expected-value calculation. Out came several trillion zeroes and a demand for somebody’s career.

The old God was gone. In his place they put the far future. It was larger, quieter and even less available for cross-examination.

Nick Bostrom built the best boxes.

Existential risk. The black ball. The singleton. Information hazards. Instrumental convergence. He has a rare gift for carving a possibility space so cleanly that everyone else moves into it and starts forwarding their mail.

But a taxonomy is not evidence. A memorable urn does not tell you how many black balls are inside it.

His careful hedging can hide how far the conclusions travel. The premise is tentative. The model is provisional. The assumptions are merely plausible. Then civilisation may need global preventive policing, or a superintelligence instructed to respect the wishes of a hypothetical cosmic host. Every sentence wears a seat belt. The vehicle is still going over the cliff.

His 2026 working paper, Optimal Timing for Superintelligence, makes the deeper problem plain. Use a person-affecting frame, count the lives of the eight billion people now alive and model the benefits of escaping ageing, and even a high catastrophe probability may justify moving quickly to AGI before pausing near deployment. The old cosmic argument made safety dominate speed. A different moral frame can make speed dominate delay.

This is not hypocrisy. Bostrom names the assumptions and follows them where they go. It is a demonstration that the policy was never written in the stars. It was balanced on the assumptions.

Toby Ord supplies the movement’s cleanest numbers: roughly one chance in six of existential catastrophe this century and one in ten from unaligned AI. The numbers travelled. Their caveats missed the flight.

Ord’s own earlier work had warned that an estimate of one in a billion is often only the probability of disaster if the argument is sound. At extreme odds, the chance that the argument itself is fucked may dominate. That warning applies to safety assurances. It applies just as hard to doom estimates.

His recent AI writing complicates the rocket story further. The Scaling Paradox argues that celebrated scaling curves often buy linear-looking gains with exponential or high-degree-polynomial resources. He published a correction when his simple hazard model for agent failures did not hold. He has not declared the risk gone. He has shown that the path is jagged and the headline number is not the model.

The movement likes his one in ten. It should learn to like the rest of his work.

The local evidence shows a community both better and worse than its caricature.

One corpus study of 49 theoretical-alignment researchers — the ILIAD analysis, a companion post to this one — covers 2,743 cleaned long-form documents across twenty-three years and 3,653 recent social posts. It finds a genuine empirical turn. Deep-learning work rose sharply. Interpretability and evaluation grew. Agent foundations declined. Roughly one substantive essay in eight contains a reversal, credited correction or null result left visible in the text.

That is admirable. Most institutions bury the scar. This one points at it.

Then a separate analysis of 728 substantive empirical-safety posts — the meta-analysis post, Nobody Is Checking — turns on the fucking lights.

About nine posts in ten name a baseline. Nine in ten state limitations. Two in three run an ablation. Yet only 19% report uncertainty, 14% report random seeds, 10% use human evaluation and 2% pre-register anything. The median design-marker score is two out of six. Not one post scores six. There are twenty-four replication posts, written by twenty-three different first authors. Exactly one person did it twice.

The field learned confession faster than verification.

Its most-read work is often its least reproducible, not mainly because famous people are lazy but because the interesting frontier models live behind corporate APIs. In the closed-model slice, code release falls from 66% in the bottom attention tier to 31% at the top. The load-bearing evidence is held by the companies the evidence is supposed to test.

The guard studies the king’s dragon with measurements supplied by the king.

These numbers need their own warning label. The corpus is forum posts, not the whole discipline. One language model extracted the claim records and no human has validated the instrument. “Ships code” means the post says code exists, not that the code runs. The honest conclusion is not that every result is rubbish. It is that the field’s standards are hard to reconcile with the size of its claims.

If your result concerns barley fertiliser, one seed is bad practice.

If your result helps decide whether governments should constrain the most powerful technology on earth, one seed is magnificent bullshit.

Then the arguments acquired an economy.

There were reading groups, fellowships and retreats. Then grants, research organisations, policy shops, university centres, career coaches, job boards and regrantors. A young person learned that AI was the most important problem. The movement offered a course explaining why. It offered a coach to redirect the young person’s career, a fellowship to provide credentials and a grant to study the claim that had brought the young person there.

This was not a conspiracy. It was an ecosystem. Ecosystems can close around a food source without anybody drawing the circle.

The biggest food source was concentrated philanthropy. Open Philanthropy, now Coefficient Giving, says it began building AI-safety institutions and talent pipelines in 2015. By 2025 it had put more than $580 million into the field, including an early $30 million grant to OpenAI.

The money funded valuable work. It also built an epistemic company town. When the same network supports researchers, training, conferences, career advice and policy organisations, dissent remains legal. It merely becomes expensive.

The parent EA community’s own 2024 survey reported respondents who were 69% men and 75% white. The survey is self-selected and the movement is broader than it. Fine. It still describes a narrow pool asked to reason on behalf of everyone.

The problem is not that white men cannot think. Many do it all fucking day. The problem is that a group can attack arguments ferociously while sharing the same priors about intelligence, progress, institutions, merit and the moral authority of clever people. Internal criticism finds errors inside the frame. It does not tell you that the frame is bent.

FTX was the ugly local test. It did not disprove effective altruism. A fraud can quote any philosopher. But senior figures were reportedly warned from 2018 onward about Sam Bankman-Fried’s dishonesty, poor controls and treatment of subordinates. The warnings were downplayed or treated as internal dispute. His money remained capable of doing enormous expected good.

People who modelled the ethics of the far future failed to process bad information about a rich man in the room.

That is not a cheap gotcha. It is a warning about the machine. If one donor may save billions of future lives, every present red flag arrives with a cosmic opportunity cost attached. Doubt becomes lost impact. The calculus need not be corrupted from outside. It can corrupt itself.

The movement set out to stop an uncontrolled intelligence race.

Its practical strategy became joining the race with nicer people.

OpenAI was built to make AGI benefit humanity. Anthropic was built by people who thought OpenAI had become unsafe. Safety researchers entered DeepMind and the frontier laboratories because that was where the models, compute, money and access were. You cannot make the frontier safe if you are not at the frontier.

That is a good argument. It is also how every weapons laboratory hires an ethicist.

The research itself is dual-use. Interpretability may reveal danger and reveal how to build a better model. Evaluations may detect hazardous capability and advertise it. Control protocols may constrain agents and become deployment infrastructure. When both safety and capability research require feeding the dragon, the accounting gets slippery.

Then the market did what markets do.

OpenAI announced a Superalignment team and promised major resources for controlling superintelligence. Less than a year later it disbanded the team; co-leader Jan Leike resigned saying safety had taken a back seat to products.

Anthropic built the clearest voluntary brake in the industry. In 2026 it removed its categorical pledge not to train or release beyond certain thresholds without adequate safeguards, replacing it with comparative commitments, road maps and risk reports. Its stated reason was competition. A unilateral stop would not help if less careful rivals kept moving.

The movement had described the unilateralist’s curse years before. Its flagship laboratory then cited the structure of the curse as the reason it could not stop.

That is the whole clusterfuck in one clean turn.

Every actor is locally rational. The funder grows the field. The researcher needs access. The safety laboratory must remain competitive. The government must not lose to another government. The advocate makes the threat vivid. The career organisation sends more people in. Each chook runs sensibly. The flock charges into the turbine.

Governments translated the old universal language into the terms states understand. The United Kingdom renamed its institute the AI Security Institute and sharpened its focus on crime, cyberattack and national security. The United States renamed its institute the Center for AI Standards and Innovation and tied the mission to competitiveness. Anthropic and OpenAI entered defence and intelligence work.

This was predictable. Governments do not have a humanity department. They have commerce departments, militaries, borders and enemies.

The 2026 International AI Safety Report gives the sober view. Fraud, cyber misuse, manipulation, unreliable agents and possible biological assistance are real or emerging. Current systems still lack the capabilities required for loss-of-control scenarios. Progress through 2030 might slow, continue or accelerate. Pre-deployment tests do not reliably predict real-world risk. Twelve companies published or updated frontier safety frameworks in 2025, but the regime remains largely voluntary.

There are serious observed harms, plausible catastrophic pathways and deep uncertainty.

The movement too often compresses that into: the god-machine is coming and we are five years from judgment.

Then they wrote the judgment as a story.

AI 2027 may be the movement’s best artifact and its worst habit in one document. Five forecasters built a concrete scenario with dates, compute estimates, fictional laboratories, Chinese espionage, recursive research and two endings. They ran war games, published their methods and invited attack. They said it was not a recommendation. They admitted prediction at that range was nearly impossible.

This is better than vague doom. A claim with a date can bleed.

The story still ate the caveats.

The title is AI 2027. The detailed timelines forecast gave all-things-considered medians of 2028, 2030 and 2033, depending on forecaster. Later model updates moved central estimates to 2029 and 2030. The authors disclosed a code bug that had pulled one median forward by about nine months.

Good. More people should admit the bug before the funeral.

But the title travels. The distribution stays home.

Now the apocalypse has a dashboard. On 29 August 2026 the live AI 2027 Tracker displayed 202 predictions, 30% evaluated and 85% accuracy. Underneath, 58 items had numerical scores. Forty-three were confirmed, nine partly accurate and five inaccurate. Roughly 70% of the full list had no verdict, and 124 items belonged to 2027.

The end of the world was pending.

Many early hits were real and mundane: agents appeared; they were expensive and unreliable; they wrote code, browsed and searched research; companies used them. But the tracker split that broad weather front into many small predictions. The first eighteen mid-2025 claims all received full marks, including several trends visible when the scenario was published. One item earned full credit for saying a run of 10²⁸ operations was roughly a thousand times GPT-4. That is arithmetic wearing a fortune-teller’s hat.

The 85% is not a calibrated probability that the scenario’s causal arc is right. It is the mean score assigned to an early, selected slice. There is no penalty for correlation, no baseline comparison, no fixed importance weighting and room for invented entities such as Agent-1 to be matched to whichever real system later resembles them.

Being right about the runway does not prove the aeroplane becomes God.

The proper use of AI 2027 is as a war game. Break it. Compare it with slower and stranger worlds. Pre-register the important claims. Keep correlated claims together. Fix resolution criteria in advance. Score the bridge from useful agents to automated research to superintelligence, not every shrub beside the road.

Otherwise the loop closes. A story creates urgency. A tracker counts the easy opening scenes. The score certifies the story. The story justifies the race. The race makes the story look prescient.

They call it forecasting. History may call it instructions.

Outside the chapel, people live in the present.

A one-month scrape of 1,337 popular AI Reddit posts — see the companion Reddit-scrape post — found five from r/ControlProblem. Five. The large stories were privacy, medical bills, jobs, model access, open weights and corporate drama. Informational posts did well. Human stakes travelled farther.

That sample is not a referendum. Reddit is not humanity. It is a thermometer showing that the people invoked by cosmic calculations often inhabit another moral weather system. They care about whether the machine helps, cheats, spies, replaces, humiliates or consoles them now.

The movement reached funders, laboratories and governments before it built a public constituency. In July 2026, 80,000 Hours made its priorities explicit. It stopped recommending mid- and senior-level roles in global health, animal welfare and climate work because transformative AI had become an “all-hands-on-deck situation.” Junior roles remained partly because they built useful career capital for later AI work.

The mosquito net had become training for the machine war.

Maybe that prioritisation is correct. That is not the same as saying it is authorised. A probability multiplied by an imagined future population does not confer democratic legitimacy. Noticing a risk early does not appoint you to represent the species. Writing “humanity” in an objective function does not mean humanity signed it.

The AI-safety movement is no longer one movement.

It is technical researchers producing evidence. It is corporate teams with access and no independence. It is independent evaluators with independence and incomplete access. It is pause activists, national-security officials, effective altruist recruiters, policy scholars, lab executives, longtermist philosophers and engineers trying to keep today’s systems from doing stupid things at scale.

The label survives because every faction benefits from it. “AI safety” sounds better than product assurance, industrial policy, civilisational triage, military reliability, corporate public relations or stopping the end of the world.

Its greatest success was making safety impossible to ignore.

Its greatest failure was allowing safety to mean everything.

When one word means extinction, discrimination, jailbreaks, battlefield reliability, child protection, national advantage, model obedience and whether a chatbot says a naughty word, the winner is whoever owns the model and the microphone.

The people in the movement are not unusually stupid. Many are unusually intelligent. That is part of the problem. Intelligence lets a person travel farther after taking the wrong road. Moral seriousness supplies the petrol. Social reinforcement removes the signs. Large numbers make the destination look inevitable.

The cure is not to sneer at catastrophic risk. The risk may be real. The cure is to take it seriously enough to remove it from the church.

Independent evaluators need legal access to frontier systems, stable public funding and whistleblower protection. Safety frameworks need external audits, incident reporting and thresholds that do not evaporate when a competitor ships. Funding networks and conflicts should be visible. Replication must become a job. Claims based on closed models must remain provisional. Forecasts must be scored by rules written before the world answers.

Present harm and catastrophic risk belong in the same analysis. Labour, civil rights, cybersecurity, public health, energy, military escalation and affected communities are not distractions from safety. They are the terrain on which any future catastrophe will be prepared or prevented.

Above all, no laboratory, philosophy seminar, philanthropic network, defence ministry or flock of clever chooks should govern the future merely because its expected-value calculation contains the largest number.

Maybe this is the hinge of history. A sane civilisation would consider that possibility.

It would also remember that every generation of ambitious bastards has believed history turned beneath its boots.

If this is the hinge, approach it with evidence, divided power, public authority and a hand on the brake.

The machine may not kill us.

The race might.

And the people paid to stop the race are already in it.


This is a polemic, not a neutral literature review. The empirical claims are linked to their sources. Local corpus studies are treated as lenses, not verdicts: the ILIAD labels are noisy; the 728-post meta-analysis uses an unvalidated single-model extraction layer; and the Reddit scrape covers one month. AI 2027 Tracker figures are a dated snapshot and will change.

LLM-to-read

Abstract

A declared polemic synthesising three companion corpus studies and public reporting. Its argument: the capital-S AI-safety movement correctly identified a real class of danger (capable systems plus bad objectives plus granted authority), then fused that insight with a status incentive — believing oneself the first to see the decisive risk — that it never learned to separate out. Consequences claimed: empirical standards that lag the size of the claims; a philanthropically concentrated ecosystem in which dissent is expensive; a practical strategy that collapsed into joining the race it meant to stop; forecasting artifacts whose caveats travel worse than their titles; and a public that inhabits a different moral register than the movement’s calculations. Prescription: move catastrophic-risk governance from movement authority to ordinary civic machinery.

Claims

Polemical theses (argued, not measured):

Factual anchors (each linked or cross-referenced in text):

Data & provenance

Method

Argumentative synthesis over the companion analyses and public reporting, written and labelled as a polemic. Empirical claims are linked to sources in-text; corpus studies are used as lenses, not verdicts.

Reproduction

Caveats

Provenance: edited September 2026.