Beware the Man of Many Studies
Two waves of study-skepticism in rationalist writing, fifteen years apart — and why the second one matters more. Part 2 of a series drawn from 37.6 million words across 34 rationalist blogs.
I. Fifty million dollars and nothing to show for it
Between 1974 and 1982 the United States government randomly gave 7,700 people medical care that was either free or not free, for up to five years, and measured who ended up healthier. The RAND Health Insurance Experiment cost fifty million dollars. The free-care group consumed 30 to 40 percent more medicine. They were not, in any measurable way, healthier.
Robin Hanson — an economist who trained in physics and helped invent prediction markets — wrote about RAND in 2008 and petitioned the government to run it again, bigger. Nothing happened. So he did what a blogger can do instead: he went through the medical literature one review at a time and reported, in three-hundred-word posts with the deadpan of a coroner, what it showed. Supplements Kill. Beware Active Placebos. Peer Review Is Random. Skip Cancer Screens: screening “consistently leads to more cancer detection and more cancer treatment, it consistently doesn’t lead to lower mortality.”
Thirteen years later he could finally post the experiment he’d petitioned for — in India, not America: 52,292 people in 435 villages randomised into hospital insurance, “very few statistically significant impacts … on health.” He is not a man who gloats. He just updates.
II. Two waves
Count everything — every post by every author on the guest list of LessOnline, the rationalist blogosphere’s annual festival, sorted into themes by a topic model — and one theme turns out to be not a subject but a practice: writing whose question is how much should I believe this study? Around 900 posts, spread across seven blogs.
It comes in two waves. The first is Hanson’s: 85 posts in 2007, then 60, 78, 68, 53 — nearly all his, nearly all about medicine. Then a trough: a dozen a year or fewer from 2014 to 2020, three in 2019. Then a second wave from different hands: 56, 51, 62, 52, 50 across 2021–2025. (One caution: the blog that carried the theme through the 2010s, Slate Star Codex, isn’t in the corpus — only its 2021 successor — so part of the trough is a missing bookshelf.)
III. The hinge
In May 2015 the physicist Steve Hsu posted 191 words on the first big replication effort in psychology — 100 findings, 39 reproduced. What he added was a diagnosis of the readers rather than the studies. Researchers “might pay lip service” to the result, he wrote, but “they typically have not updated their posteriors to reflect the low reliability of research results, even in the top journals.”
That is the hinge between the waves. The first asked does the medicine work? — a question about the world. The second asks what should a reader do with a literature this unreliable? — a question about the reader, which turns out to be much harder.
IV. The second wave argues with itself
Its founding document is a rebuttal. In June 2023 the pseudonymous Cremieux published Beware the Man of Many Studies — 4,000 words that begin by praising Scott Alexander’s famous 2014 advice (don’t trust one study; read the whole literature) and then turn on it. The subtitle is the argument: “Low-quality literatures mean meta-analyses are frequently worse than single studies.” A hundred bad studies averaged together do not become one good study. They become a confident wrong answer with a narrow confidence interval.
Scott, meanwhile, was policing the other flank. Against That Poverty And Infant EEGs Study (2022) dismantles a paper the Times had celebrated as “basically a typical null-result-having study.” But All Medications Are Insignificant In The Eyes Of God And Traditional Effect Size Criteria (2023) argues that the fashionable “effect size 0.30 is meaningless” critique of antidepressants would also condemn drugs that obviously work — so “downweight all claims about ‘this drug has a meaningless effect size’ compared to your other sources of evidence, like your clinical experience.”
Notice what happened. In the first wave the skeptic was one economist reading against the grain. In the second, skepticism is the default — so the interesting posts are the ones policing skepticism itself. Cremieux: the meta-analysis you were told to trust can be worse than the study you were told to distrust. Scott: the heuristic you were told to apply will tell you Ambien doesn’t work. Skepticism has become a tool that needs its own calibration.
V. Why this is the pressing one
The community that built this discipline now spends most of its words on artificial intelligence, and the AI argument is conducted almost entirely in this theme’s currency. Timelines are forecasts with error bars. Capabilities are benchmark scores published by the labs being benchmarked. Safety claims are evaluation results with sample sizes and funders. Every failure mode this theme catalogued in medicine is present in the AI evidence base, at higher stakes and faster tempo.
Hanson petitioned for an experiment in 2008 and got one in 2021. That is how long good evidence on a hard question takes. The AI question does not have thirteen years. So a skill learned one Cochrane review at a time, by people who did not know what they were training for, is now the load-bearing skill of the community’s central argument. And Hsu’s line is the whole theme in a sentence: the hard part was never showing the studies were unreliable. The hard part is that people “have not updated their posteriors” — and updating is the one thing this community claims, above all else, to know how to do.
LLM-to-read
-
Abstract — In a corpus of 34 rationalist blogs (LessOnline attending authors; ~21.4k articles, 37.6M words, 2005–August 2026), “Empirical studies & medical evidence” — writing whose subject is how much to trust research evidence, and whether medicine works — is a mid-sized theme (898 articles, 7th by count, 2.1M words) with wide spread (7 of 34 blogs ≥5% of posts). It arrives in two waves with different authors, subjects, and stances: a 2007–2011 single-author campaign against the medical literature, and a 2021– community-wide, self-policing discipline of evidence-weighting that now underwrites the community’s AI argument.
-
Claims
- Theme size: 898 articles, 2.1M words, 7th by article count; 7 of 34 blogs devote ≥5% of posts to it.
- Articles per year — Wave 1: 2007: 85, 2008: 60, 2009: 78, 2010: 68, 2011: 53 (almost entirely Robin Hanson, Overcoming Bias, medical-literature skepticism). Trough 2014–2020: 3–12/yr. Wave 2: 2021: 56, 2022: 51, 2023: 62, 2024: 52, 2025: 50, 2026 through August: 27 (ACX, Cremieux, gwern, Constantin, Aella).
- By blog: Overcoming Bias 400, ACX 134, gwern 84, Cremieux 76, infoproc (Hsu) 54, Constantin 36, Friedman 32, Aella 19.
- Wave 1 asks “does the medicine work?”: Hanson’s run — Random Smoking Trials (2009), Supplements Kill, Beware Active Placebos, Peer Review Is Random (2010), Skip Cancer Screens (2011) — anchored in the RAND Health Insurance Experiment (1974–82, n=7,700, ~$50M; free care raised consumption 30–40% with no measurable health effect) and closed by the Karnataka Hospital Insurance Experiment (2021, n=52,292, 435 villages, same null).
- Hinge: Hsu, Replication is hard; understanding what that means is even harder (2015-05-03, 191 words; 100 findings, 39 reproduced) — the problem is readers not updating, not just studies not replicating.
- Wave 2 asks “what should a reader do with an unreliable literature?” and self-corrects: Cremieux, Beware the Man of Many Studies (2023-06-05) argues meta-analyses of low-quality literatures are often worse than single studies, rebutting SSC’s 2014 Beware the Man of One Study; ACX, Against That Poverty And Infant EEGs Study (2022-01-26) and All Medications Are Insignificant In The Eyes Of God And Traditional Effect Size Criteria (2023-05-31) police both credulity and over-skepticism (the 0.30 effect-size heuristic would condemn drugs that work).
- Strand 3, generate your own evidence: gwern’s How Often Does Correlation=Causality? compilation (19k words; correlation ≠ causality “a large fraction of the time, although probably not a majority”) plus self-directed evidence (Creatine Cognition Meta-analysis, Bacopa Quasi-Experiment, Lunar circadian rhythms on his own sleep data); Aella’s large-n surveys (from The Psychedelic Intensity Scale, 2018); Constantin’s trial-mining (Progesterone for Postpartum Depression, 2022).
- Why pressing: the community’s dominant current theme (AI; see part 1) is argued in this theme’s currency — forecasts, benchmarks, lab-published evals — reproducing every catalogued failure mode (single dramatic study, low-quality literature averaged into false confidence, misapplied heuristics, non-updating readers) at higher stakes and shorter timescales.
-
Data & provenance — Every retrievable post by every author on the LessOnline festival guest lists: 34 blogs, ~21.4k articles (21,403 at snapshot), 37.6M words, 2005–August 2026; obtained by scraping the blogs’ public archives; corpus snapshot 2026-08-27.
-
Method — NMF topic decomposition (K=40) over TF-IDF of all articles ≥150 words, components hand-labelled into 23 themes. This theme merges two components — research methodology (results, data, study, sample, studies, estimate, average, effect, bias, survey, significant) and medicine (health, medical, patients, medicine, doctors, treatment, drug, care, insurance, trials, clinical). Articles assigned if the theme carries ≥15% of NMF mass.
-
Reproduction — The post names no commands; per the series’ shared pipeline:
theme_analysis.py(K=40) →summaries/_themes.json; per-blog article data inscraped/<slug>/articles.jsonl. -
Caveats — Slate Star Codex (2013–2020) is absent from the corpus (only ACX 2021– is present), so part of the 2014–2020 trough is a coverage gap, not a real lull. Overcoming Bias’s 2007–2008 posts include ~295 by Yudkowsky (then a co-blogger); a handful of Wave 1 posts are his rather than Hanson’s. Theme assignment is statistical; individual mislabels exist, aggregates are robust.
-
Provenance — Edited September 2026.