The Man in the Lab Coat
Part 2 of a three-part series from the same corpus.
On a corpus where one report in five is written as a timestamped log — what the apparatus of rigour buys, and what it quietly doesn’t.
I. An apple, and a control group
In September 2010 a man took LSD in a shared house and went downstairs for a snack.
I donned my white labcoat… Just wearing it provokes me to think within a scientific framework. Suddenly, Im not just getting a snack from the kitchen. Now I am performing a perceptual test… I wander around the house, feeding slices of apple to people. Control group: People not on LSD! How would you describe this apple?
The coat is wonderful — that he owns one, that he admits to using a costume as a cognitive prosthetic. But the next line matters more:
My control group told me is was sour, and crisp. Those subjects on LSD were in agreement.
Sober people: sour and crisp. Tripping people: sour and crisp. He ran the comparison, got a null, wrote it down, and moved on. Nobody made him. He was in a house full of drunks at two in the morning holding a plate of apple slices.
And my classifier scored that report negative for every scientific-register theme I built. The most rigorous act in the corpus is invisible to the instrument built to find rigour. I’ll come back to that.
II. More instrument, less sermon
I categorised 25,171 experience reports — 26.8 million words across four decades. The top of the list is what you would guess: nausea, closed-eye geometry, the world breathing, time coming apart.
Third is not a topic at all. It is a format. Roughly one report in five is written as a timestamped log — a document whose spine is a column of time-marks rather than sentences.
The classifier catches these at 93% precision — the cleanest signal I measured, cleaner than nausea, because prose almost never begins a line with T+2:30.
And the form is winning. Standardising for report length, the timestamped log rose about 60% in relative terms between 2000 and 2026, while the advice-to-the-reader coda — be careful, respect the substance, always have a sitter — fell 47%. Those are the only two era trends that survived length-standardisation; four others were nothing but word count. In twenty-five years the reports stopped preaching at you and started handing you a dataset.
III. What the timestamp buys
The property that made the classifier work is the point: the timestamp is legible to a machine. “After a while things got weird” is a story; a column of time-marks is data, and it is data because the author accepted a constraint that cost him something in the moment.
From a 2018 LSD report, “Famous Last Words”:
I set a bunch of alarms on my phone, one every 30 minutes for 10 hours, to remind me to take note of what I was experiencing.
Twenty alarms, on LSD. The entries record not visionary content but calibration:
T: 0hr30m — I feel a little light. I’m probably imagining it. T: 0hr:44m — Not imagining it. Real. I like it.
Fourteen minutes. He registers an effect, discounts it as expectancy, then revises. That is live placebo discrimination on oneself, timestamped — and only the format preserves it. Written as prose the next morning it becomes “it came on slowly.” The clock keeps the epistemics.
IV. And what it doesn’t
I ran a recall audit — thirty reports the classifier had rejected, read in full — and asked a broader question of each: does this narrator deliberately record or measure in any form?
Seventeen of thirty. Roughly four times the rate of the timestamp format itself. Cassette recorders. Digital scales against a remembered reference dose. Volumetric dosing — 265mg in 500ml, 50ml decanted — with a note flagging a contaminant as a confound. The timestamped log is not the phenomenon; it is the visible tip. Treating your own experience as an experiment is close to a majority practice, and the format is merely the part a regex can see.
And yet. Almost nobody has a comparison condition — the lab-coat man is remarkable precisely because he is nearly alone. Nobody is blinded. The dose stays a guess where it matters most: null rates track dose uncertainty, 15.7% for wild-picked Amanita against 5.1% for nitrous. Timestamps do not tell you what was on the blotter.
Most sharply: the instrumentation does not prevent the escalation. Reports containing a null are four times more likely to contain a redose, and that is no lower among the meticulous. The man with twenty alarms is doing something real. He is not thereby protected. The apparatus records the decision; it does not improve it.
V. Cheap rigour and expensive rigour
This is not a story about drugs. It is about what a community gets when it adopts the forms of science without the institutions — and the honest answer has two halves people want to collapse into one.
Rationalist culture is unusually invested in visible apparatus: epistemic-status headers, confidence intervals on casual claims, calibration curves, Brier scores. People wearing lab coats because the coat provokes them to think within a scientific framework. This mostly works. It is not cargo cult: twenty-five years of drift toward the timestamp and away from the sermon is a community teaching itself something true.
The failure mode is that the apparatus is cheap and the substance is expensive, and they feel identical from the inside.
A timestamp costs nothing. A confidence interval costs a keystroke. A control group costs you a friend, an apple, and the willingness to find out that the sober people and the tripping people said the same word. Blinding costs a collaborator and your own certainty. Pre-registration costs the option of deciding afterwards what you were testing.
Every cheap thing produces the feeling of rigour; only the expensive ones produce the thing. And because the cheap ones are what is visible in a document, they are what gets rewarded and imitated.
So: which of your epistemic habits would survive if nobody could see them? The timestamp is legible, and that is most of why it spread. The control group in the kitchen at two in the morning was legible to nobody, produced a null, and was written down anyway.
VI. My own lab coat
I built 356 regular expressions, ran them across 26.8 million words, and produced a beautifully ranked table of the ten most common themes in the largest archive of first-person drug experience there is.
Then I had the matches read by hand, and roughly half the constructs turned out to be measuring something else. bad trip indicated genuine terror one time in eleven — a topic label, appearing mostly inside “I have never had a bad trip.” My detector for lye, a real extraction reagent, fired on people who lye down on the bed, and on the filmmaker Len Lye.
The table had the form of a result: columns, percentages, a ranking. Every statistical check passed. A bad instrument does all of that perfectly well.
What exposed it was 670 matches read by a person, one at a time, asking of each: is this actually what I said it was? That is the expensive thing. It produced no new numbers — it only corrected old ones, downward.
I had spent a day in a very good lab coat.
The apple is the hard part. It always was.
Corpus: 25,171 reports — 24,724 from Erowid’s Experience Vaults, 447 from PsychonautWiki (CC BY-SA 4.0, © PsychonautWiki contributors) — 26.8 million words. Erowid’s terms prohibit bulk download and AI-type analysis without written permission; the scrape proceeded on a stated permission that could not be independently verified, the project’s load-bearing and non-technical assumption. PsychonautWiki’s operators note that contributors did not consent to AI-training use. Quotations are short excerpts from individual reports, retained for verifiability.
LLM-to-read
-
Abstract — In a corpus of 25,171 first-person drug experience reports (26.8M words, 1968–2026), the third most common theme is not an experience but a format: roughly one report in five is written as a timestamped log, and over twenty-five years the corpus drifted toward that format (+~63% length-standardised) while advice-to-the-reader codas fell (−47%). A recall audit finds deliberate self-measurement in 57% of classifier-negative reports — near-majority practice — yet the apparatus coexists with no controls, no blinding, no dose certainty, and no protection against post-null redosing. The post’s argument: the corpus is a natural experiment in what the forms of rigour achieve without institutions — cheap rigour (timestamps, headers) and expensive rigour (controls, blinding, checking your own instrument) feel identical from the inside, and only the cheap kind is visible, so it is what spreads. The post applies the same critique to its own 356-regex instrument, roughly half of whose constructs were measuring something else.
-
Claims
- Timestamped-log format: 11.0% raw document frequency at 93% precision — the highest-precision construct of the 38 built (line-anchored patterns prose rarely produces). Recall audit: 4 misses in 30 negatives → true prevalence ≈ 22%, recall ≈ 46% (“one in five”).
- Broader “instrumented” construct: 17 of 30 classifier-negative reports (57%) show deliberate recording or measuring — weighed or volumetric dosing, real-time notes, cassette/voice recording, comparison against a reference dose, replication instructions, running diaries — ≈4× the timestamp rate. Single rater, broad definition, one 30-item sample: directional, not exact.
- Length-standardised era trends (observed/expected), the only two of six to survive: timestamped log 0.73 (2000–04) → 1.19 (2020–26), ≈+63% relative; advice-to-reader coda 1.43 → 0.76, ≈−47% relative. A constant false-positive rate cancels in the ratio, so the trends are robust to absolute precision. Related constructs: preparation/weighing/extraction 7.7% raw at 60% precision (leaked on
lyematching “lye down” and Len Lye); read-up-beforehand 3.2%. - What instrumentation delivers: machine legibility → comparability (the 93% precision is itself the finding); preserved in-the-moment epistemics — E111357 (“Famous Last Words”, 20 phone alarms over 10 hours) records live placebo discrimination (
T+0h30m“I feel a little light. I’m probably imagining it.” →T+0h44m“Not imagining it. Real.”); and genuine within-subject design achievable solo — E52809: reference run vs trial run, stated hypothesis, reproducible prep, BP 144/84 at a named timepoint, in 232 words, behind an “AFOAF” legal shield. - What it does not deliver: comparison conditions are near-absent across 25,171 reports; blinding is zero; dose stays uncertain (null rate 15.7% Amanita / 14.7% HBWR / 13.9% cacti / 13.0% blotter LSD vs 5.1% nitrous / 6.8% DMT / 7.3% ketamine); and no protection against escalation — P(redose | null) = 4.60% vs 1.15% baseline, RR 4.0×, not abolished by meticulousness. The apparatus records the decision; it does not improve it.
- Transfer: cheap rigour (timestamps, epistemic-status headers, stated confidence intervals, Brier scores on self-selected questions, spreadsheets) produces the feeling; expensive rigour (a comparison condition, blinding, pre-registration, dose control, publishing nulls, hand-checking your own instrument) produces the thing; only the cheap kind is visible in a document, so it is what communities reward and imitate. Diagnostic: which epistemic habits would survive unobserved? The forms are not worthless — the 25-year sermon→schema drift is a real improvement; form is necessary and radically insufficient.
- Self-application (load-bearing): 356 regexes over 26.8M words produced a ranked top-10 table; hand-adjudication of 670 matches found ~half the constructs measuring something else —
bad tripindicated genuine terror 1 time in 11 (a topic label, mostly inside “I have never had a bad trip”); bare “went outside” carried 67.8% of nature hits and usually meant a cigarette. Every statistical check passed on the bad version. Consistency is not validity. - Recall point: the instrument finds typical instances, not exemplary ones — of three hand-picked exemplars of the scientific register it caught one, and E80897 (lab coat, apple, control group, recorded null: the most rigorous act in the corpus) matches nothing. Reading finds the best cases; counting finds the common ones; neither substitutes for the other.
- Key reports: E80897 (lab coat/control group) · E111357 (“Famous Last Words”) · E52809 (reference vs trial run, BP readings, AFOAF) · E42497 (notes written so “my sober self” would know) · E13852 (pre-committed note to future self) · E73055 (null filed as a product review).
-
Data & provenance — 25,171 first-person drug experience reports, 26.8M words, 1968–2026: 24,724 Erowid Experience Vaults (deduplicated) + 447 PsychonautWiki. Editorial interpolations stripped and verified (0/60 residual by audit). Obtained by the scrape documented in the predecessor post: past Cloudflare, against Erowid’s stated terms (no download, analysis or AI-type use without written permission), on a stated permission that could not be independently verified; PsychonautWiki content is CC BY-SA 4.0 and its contributors did not consent to AI-training use.
-
Method — 38 regex-lexicon themes (356 patterns) over the full corpus; per-theme precision audited on 25–30-item match samples and recall on 30 classifier negatives read in full; era trends length-standardised by indirect standardisation on word-count deciles (observed/expected); 670 matches hand-adjudicated across the instrument; a broader “does this narrator deliberately record or measure?” question asked of the recall sample.
-
Reproduction — No commands or paths published in the post. Reproduction requires the project’s lexicon and the corpus; the corpus cannot be redistributed under Erowid’s terms.
-
Caveats — Precision from 25–30-item samples (±~10pp, 1 SE); recall from 30 negatives with a single adjudicator per theme; no inter-rater reliability; the 57% “instrumented” figure rests on one broad-definition sample and needs replication; known residual — right-side negation (“the wood grain lacks breathing”) escapes the left-window veto, affecting ~2% of visual-theme hits; the corpus samples drug writing, not drug use. §IV’s “no lower among the meticulous” is stronger than this section’s “not abolished by meticulousness” — both left unreconciled; the era-trend delta is stated here as ≈+63% to match Part 3’s figure for the identical trajectory (0.73 → 1.19), and §II’s “about 60%” is a rounding of the same figure.
-
Provenance: edited September 2026.