Nothing Happened
Part 1 of a three-part series from the same corpus.
What the world’s largest archive of uncontrolled self-experimentation does with its failures.
I. Four hypotheses
In September 2001 someone posted 330 words to Erowid titled “Possible Immunity to Ketamine?” He took what he was told was ketamine, waited, and got almost nothing — “Basically, it sucked.” Then:
Possible conclusions: A) I am immune. B) It wasn’t K. C) It was not enough K for me. D) I have to break my tolerance for it.
He declines to pick one, because he has no way to distinguish them. But look at the shape of the list: three of the four recommend taking more. Only B — it wasn’t the drug — concludes the substance doesn’t work. That asymmetry runs through everything below.
II. One report in six
I categorised 25,171 experience reports — 26.8 million words, four decades. 320 were independently hand-coded by eight readers, a lexicon covered the rest, and every theme’s matches were then hand-audited, because the previous study on this corpus — Twenty-four thousand trip reports — found its headline construct was three-quarters noise.
Somewhere between one report in seven and one in five is about the drug not working. The corrected instrument says 13.6%; the eight independent readers say 20.0%. Call it one in six.
Nobody has to file these. No journal, no grant renewal, no reviewer asking for negative controls, and the genre’s entire appeal is the extraordinary. One in six says: I did the thing, and the thing did not happen.
“Ghana Strain Review” (2008), 255 words: a man buys the cheap, weaker strain of Hawaiian Baby Woodrose seeds, reasons that he will compensate with dose, eats a hundred and twenty-five of them over two hours, and gets a slight buzz. He files it anyway, and says why in his first line:
I noticed that there aren’t many (if any) HBWR Ghana strain reports so I thought I should add this as it is happening.
He noticed a gap in the literature and filled it with a null.
III. The dud rate measures dose uncertainty
The nulls are not spread evenly, and the ordering is almost too neat.
At the top: Amanita muscaria 15.7%. Hawaiian Baby Woodrose 14.7%. Columnar cacti 13.9%. Datura 13.1%. Blotter LSD 13.0%. At the bottom: nitrous oxide 5.1%. DMT 6.8%. Ketamine 7.3%.
The top is wild-picked botanicals and blotter you cannot assay by looking at it. The bottom is material you inhale or inject, where onset arrives in seconds. This is not a fact about pharmacology. It is a measurement of how uncertain the dose was. Nitrous sits at 5% because ten seconds after the balloon you know. Amanita sits at 16% because you never know.
IV. The forty-five minute problem
My classifier for this theme was the worst of the thirty-eight I built: 44% precision, 34% recall. The adjudicator’s list of its errors was nearly all one error — “nothing happened but 10 minutes later I started trippin”, “nothing for an hour, then the acid hit”.
The machine could not tell “nothing happened” from “nothing has happened yet.”
Neither can the person holding the empty cup. That is not a defect of my regex; it is the phenomenology. At T+45 minutes, “this was a dud” and “this hasn’t started” are the same sentence — and they come apart only after you have had to decide what to do.
An AET report from 2007, logged at T+2:30:
In fact, we aren’t even sure we could distinguish this feeling from placebo.
A man mid-experiment, correctly identifying that his instrument — himself — cannot resolve the signal from zero. The next entry is timestamped T+3:15: he retrieves Dilaudid he had dissolved and micron-filtered a month earlier, draws up half a cc, and injects it.
Forty-five minutes from that recognition to the syringe. Corpus-wide, reports containing a null are four times more likely to also contain a redose — 4.60% against 1.15%, a ratio that survived a full rebuild of the classifier.
The uninterpretable null is the dangerous moment — not the big dose, which is usually decided in advance by someone who has read something. It is the moment where evidence is absent rather than negative, and absence must be converted into an action by someone who has already spent the money and the evening.
V. Why this is a rationalist problem
The drugs are mostly a distraction.
Rationalist culture runs on n=1 self-experimentation — nootropics, microdosing, sleep interventions, supplements, protocols — which is genuinely admirable: people try things on themselves and report back rather than waiting twenty years for an underpowered RCT.
The reporting is also catastrophically biased in a way Erowid’s is not. The post gets written when something worked. Nobody writes “six weeks, indistinguishable from baseline” — not from dishonesty, but because there is no story there, and because a null feels like it reflects on you: wrong dose, wrong brand, not enough discipline. Three of four hypotheses point at trying harder.
Erowid gets this right for a structural reason worth stealing: the archive is indexed by substance, not by outcome. You do not file because your experience was noteworthy; you file under the seeds’ page because that is where reports about those seeds go, and the row exists whether or not anything happened. The dud reports are not heroic. They are product reviews. Someone left one star on a plant, and the corpus got better.
Self-experimentation communities index by outcome — the post exists because the result was interesting — and so they lose their negatives. That is the whole difference, and it is a filing convention, not a virtue.
The corollary is the 5%-to-16% gradient: a null is uninterpretable unless you controlled the dose. Unknown bioavailability, unknown adherence and a subjective endpoint is a self-built Amanita. You will generate nulls you cannot read, and the only move the situation offers is more.
VI. A null of my own
I should turn this on the analysis, or I am doing the thing I am complaining about. My headline theme was measured by my worst tool, and I only know it because 250 of its matches were read by hand. Every statistical check had passed: prevalence stable across eras once standardised for length, replicated across two independent corpora, sensibly correlated with substance. A bad instrument does all of that perfectly well.
There is a version of this post built on the raw number — “14.4% of trip reports describe the drug failing” — that is clean, confident, and wrong in both directions at once. The honest version is somewhere between one in seven and one in five, and here is why I can’t do better.
The reason I can’t is that “it didn’t work” has no canonical phrasing, no agreed boundary against “it worked slowly,” and no fact of the matter at the moment it is being lived. The category is hard to measure for exactly the reason it is hard to live through.
Corpus: 25,171 reports — 24,724 from Erowid’s Experience Vaults, 447 from PsychonautWiki (CC BY-SA 4.0, © PsychonautWiki contributors) — 26.8 million words. Erowid’s terms prohibit bulk download and AI-type analysis without written permission; the scrape proceeded on a stated permission that could not be independently verified, the project’s load-bearing and non-technical assumption. PsychonautWiki’s operators note that contributors did not consent to AI-training use. Quotations are short excerpts from individual reports, retained for verifiability.
LLM-to-read
-
Abstract — A categorisation of 25,171 first-person drug experience reports (26.8M words) finds that roughly one report in six documents the substance producing markedly less than expected, or nothing — a large, voluntarily published body of negative results in an uncompensated, legally exposed literature. The per-substance null rate tracks dose uncertainty rather than pharmacology (unassayable botanicals and blotter at the top, instant-onset inhaled or injected material at the bottom), and reports containing a null are four times more likely to also contain a redose, because at T+45 minutes “dud” and “not yet” are indistinguishable from the inside. The theme’s own classifier was the worst of the 38 built, and every statistical check passed on it anyway; only hand-reading exposed the error rates. The transfer claim: self-experimentation communities lose their nulls because they index by outcome, where Erowid indexes by substance.
-
Claims
- Central: roughly 1 in 6 reports (range 14–20%) is about the drug not working. Two independent estimates: the instrument with both error rates corrected gives 13.6% (10.54% raw × 44% precision, plus a 10% miss rate measured on 30 audited negatives); eight hand-coders across 320 full reports give 20.0% pooled on the decks that coded the theme (8/40, 10/40, 11/40, 8/40, 3/40). The spread is caused by boundary ambiguity, not sampling — do not report a point estimate.
- The dud rate measures dose uncertainty, not pharmacology. High: Amanita muscaria 15.7% · Hawaiian Baby Woodrose 14.7% · columnar cacti 13.9% · Datura 13.1% · blotter LSD 13.0% · morning glory 12.6%. Low: nitrous oxide 5.1% · DPT 6.4% · DMT 6.8% · ketamine 7.3% · MXE 7.5%. Corpus baseline 10.5% raw. Ordering principle: unassayable dose + ambiguous onset → high; fixed-dose inhaled/injected route + instant onset → low.
- The null causes escalation: P(redose | null present) = 4.60% vs P(redose | no null) = 1.15%, RR 4.0×; stable across a full classifier rebuild (v1 RR 4.35×), so not an artefact of one lexicon. Mechanism: the classifier’s dominant error — confusing “nothing happened” with “nothing has happened yet” — is also the subject’s real-time epistemic situation, resolved into action under sunk cost (E61822: cannot distinguish from placebo at T+2:30; injects Dilaudid at T+3:15).
- The instrument is the worst of the 38 built:
dud_null_resultat 44% precision, 34% recall (compareclock_logat 93% precision). Causes: no canonical phrasing for absence; “disappointed” attaches to non-drug objects; come-up impatience is lexically identical to a genuine null; missed instances are graded/backgrounded rather than absent-framed. - Methodological: every statistical check passed on the bad instrument — stable across eras after length standardisation, replicated across two corpora, sensible per-substance correlations. Only hand-reading 250 matches exposed it. Consistency is not validity.
- Transfer: (1) a null is uninterpretable without dose control — the 5%→16% gradient is the lesson; unknown bioavailability + unknown adherence + subjective endpoint is a self-built Amanita; (2) pre-commit the post-null action, not just the pre-dose decision — the badly made decision happens at T+45 under sunk cost with evidence absent rather than negative (corpus example of the fix: E13852’s written note to a future compromised self); (3) publish nulls by indexing archives by substance, not by outcome — the slot exists whether or not anything happened, so a null is a routine filing rather than an admission; (4) the four-hypothesis pattern (E7402: immune / wasn’t K / not enough / tolerance) has three of four branches recommending escalation — a structural asymmetry in how self-experimenters explain their own nulls.
- Key reports: E7402 (hypothesis enumeration) · E73055 (“Ghana Strain Review”: null filed to fill a literature gap) · E61822 (placebo-indistinguishable → injects Dilaudid) · E2075 (“Try A Lower Dose”: programmer escalates to six packs) · E13852 (pre-committed note to future self) · E109009 (escalating nulls).
-
Data & provenance — 25,171 first-person drug experience reports, 26.8M words, 1968–2026: 24,724 Erowid Experience Vaults reports (deduplicated) + 447 PsychonautWiki reports. Median report 821 words. Editorial interpolations stripped (1,941 Erowid curator notes, 392 PsychonautWiki “Effects analysis” blocks); residual contamination 0/60 by audit. Obtained by the scrape documented in the predecessor post: past Cloudflare, against Erowid’s stated terms (no download, analysis or AI-type use without written permission), on a stated permission that could not be independently verified; PsychonautWiki content is CC BY-SA 4.0 and its contributors did not consent to AI-training use.
-
Method — The series’ 38-theme regex lexicon (356 patterns) run over all 25,171 reports; 320 reports independently hand-coded by eight readers; per-theme precision audited on samples of matches (250 read for this theme) and recall on 30 classifier negatives; prevalence length-standardised on word-count deciles; the classifier fully rebuilt once as a stability check on the redose ratio.
-
Reproduction — No commands or paths published in the post. Reproduction requires the project’s lexicon and the corpus; the corpus cannot be redistributed under Erowid’s terms.
-
Caveats — Precision from 25-item samples (±~10pp, 1 SE); recall from 30 negatives with a single adjudicator; no inter-rater reliability on the audit; Erowid’s population is self-selected and pseudonymous; substance identity is usually unverified (itself one of the four hypotheses); the corpus samples drug writing, not drug use. §VI’s “14.4%” matches neither the stated 10.5% raw nor the 13.6% corrected figure (possibly a v1-classifier number) — left unreconciled; the per-substance rates differ from the predecessor study’s prior-dud rates (Amanita 21.7% there vs 15.7% here) — different construct and instrument, left unreconciled in both.
-
Provenance: edited September 2026.