Foomax

Nothing Happened

September 2026 · Erowid

Part 1 of a three-part series from the same corpus.

What the world’s largest archive of uncontrolled self-experimentation does with its failures.

I. Four hypotheses

In September 2001 someone posted 330 words to Erowid titled “Possible Immunity to Ketamine?” He took what he was told was ketamine, waited, and got almost nothing — “Basically, it sucked.” Then:

Possible conclusions: A) I am immune. B) It wasn’t K. C) It was not enough K for me. D) I have to break my tolerance for it.

He declines to pick one, because he has no way to distinguish them. But look at the shape of the list: three of the four recommend taking more. Only B — it wasn’t the drug — concludes the substance doesn’t work. That asymmetry runs through everything below.

II. One report in six

I categorised 25,171 experience reports — 26.8 million words, four decades. 320 were independently hand-coded by eight readers, a lexicon covered the rest, and every theme’s matches were then hand-audited, because the previous study on this corpus — Twenty-four thousand trip reports — found its headline construct was three-quarters noise.

Somewhere between one report in seven and one in five is about the drug not working. The corrected instrument says 13.6%; the eight independent readers say 20.0%. Call it one in six.

Nobody has to file these. No journal, no grant renewal, no reviewer asking for negative controls, and the genre’s entire appeal is the extraordinary. One in six says: I did the thing, and the thing did not happen.

“Ghana Strain Review” (2008), 255 words: a man buys the cheap, weaker strain of Hawaiian Baby Woodrose seeds, reasons that he will compensate with dose, eats a hundred and twenty-five of them over two hours, and gets a slight buzz. He files it anyway, and says why in his first line:

I noticed that there aren’t many (if any) HBWR Ghana strain reports so I thought I should add this as it is happening.

He noticed a gap in the literature and filled it with a null.

III. The dud rate measures dose uncertainty

The nulls are not spread evenly, and the ordering is almost too neat.

At the top: Amanita muscaria 15.7%. Hawaiian Baby Woodrose 14.7%. Columnar cacti 13.9%. Datura 13.1%. Blotter LSD 13.0%. At the bottom: nitrous oxide 5.1%. DMT 6.8%. Ketamine 7.3%.

The top is wild-picked botanicals and blotter you cannot assay by looking at it. The bottom is material you inhale or inject, where onset arrives in seconds. This is not a fact about pharmacology. It is a measurement of how uncertain the dose was. Nitrous sits at 5% because ten seconds after the balloon you know. Amanita sits at 16% because you never know.

IV. The forty-five minute problem

My classifier for this theme was the worst of the thirty-eight I built: 44% precision, 34% recall. The adjudicator’s list of its errors was nearly all one error — “nothing happened but 10 minutes later I started trippin”, “nothing for an hour, then the acid hit”.

The machine could not tell “nothing happened” from “nothing has happened yet.”

Neither can the person holding the empty cup. That is not a defect of my regex; it is the phenomenology. At T+45 minutes, “this was a dud” and “this hasn’t started” are the same sentence — and they come apart only after you have had to decide what to do.

An AET report from 2007, logged at T+2:30:

In fact, we aren’t even sure we could distinguish this feeling from placebo.

A man mid-experiment, correctly identifying that his instrument — himself — cannot resolve the signal from zero. The next entry is timestamped T+3:15: he retrieves Dilaudid he had dissolved and micron-filtered a month earlier, draws up half a cc, and injects it.

Forty-five minutes from that recognition to the syringe. Corpus-wide, reports containing a null are four times more likely to also contain a redose — 4.60% against 1.15%, a ratio that survived a full rebuild of the classifier.

The uninterpretable null is the dangerous moment — not the big dose, which is usually decided in advance by someone who has read something. It is the moment where evidence is absent rather than negative, and absence must be converted into an action by someone who has already spent the money and the evening.

V. Why this is a rationalist problem

The drugs are mostly a distraction.

Rationalist culture runs on n=1 self-experimentation — nootropics, microdosing, sleep interventions, supplements, protocols — which is genuinely admirable: people try things on themselves and report back rather than waiting twenty years for an underpowered RCT.

The reporting is also catastrophically biased in a way Erowid’s is not. The post gets written when something worked. Nobody writes “six weeks, indistinguishable from baseline” — not from dishonesty, but because there is no story there, and because a null feels like it reflects on you: wrong dose, wrong brand, not enough discipline. Three of four hypotheses point at trying harder.

Erowid gets this right for a structural reason worth stealing: the archive is indexed by substance, not by outcome. You do not file because your experience was noteworthy; you file under the seeds’ page because that is where reports about those seeds go, and the row exists whether or not anything happened. The dud reports are not heroic. They are product reviews. Someone left one star on a plant, and the corpus got better.

Self-experimentation communities index by outcome — the post exists because the result was interesting — and so they lose their negatives. That is the whole difference, and it is a filing convention, not a virtue.

The corollary is the 5%-to-16% gradient: a null is uninterpretable unless you controlled the dose. Unknown bioavailability, unknown adherence and a subjective endpoint is a self-built Amanita. You will generate nulls you cannot read, and the only move the situation offers is more.

VI. A null of my own

I should turn this on the analysis, or I am doing the thing I am complaining about. My headline theme was measured by my worst tool, and I only know it because 250 of its matches were read by hand. Every statistical check had passed: prevalence stable across eras once standardised for length, replicated across two independent corpora, sensibly correlated with substance. A bad instrument does all of that perfectly well.

There is a version of this post built on the raw number — “14.4% of trip reports describe the drug failing” — that is clean, confident, and wrong in both directions at once. The honest version is somewhere between one in seven and one in five, and here is why I can’t do better.

The reason I can’t is that “it didn’t work” has no canonical phrasing, no agreed boundary against “it worked slowly,” and no fact of the matter at the moment it is being lived. The category is hard to measure for exactly the reason it is hard to live through.


Corpus: 25,171 reports — 24,724 from Erowid’s Experience Vaults, 447 from PsychonautWiki (CC BY-SA 4.0, © PsychonautWiki contributors) — 26.8 million words. Erowid’s terms prohibit bulk download and AI-type analysis without written permission; the scrape proceeded on a stated permission that could not be independently verified, the project’s load-bearing and non-technical assumption. PsychonautWiki’s operators note that contributors did not consent to AI-training use. Quotations are short excerpts from individual reports, retained for verifiability.

LLM-to-read