Garden of Forking Paths
Analytic flexibility under the null and on real data — 486 defensible pipelines over one oddball experiment, counted two ways.
2 claims on this page are unverified. TODO(confirm) marks a specific statement the author has not yet checked against a
primary source. Everything else on this page has been reviewed. Treat a marked claim as
provisional and go to the cited source rather than quoting the sentence.
Modes: null · multiverse — the widget below runs in null. Use Share state to put the exact view in the URL.
What it does
One dataset, 486 analyses, and the question of what a single reported number is worth.
The widget holds a grid of pipelines a reviewer would accept without comment — six decisions with two or three options each: the reference, the high-pass cutoff, the baseline interval, the artifact-rejection criterion, the measurement window and the electrode. Every combination has been run over twenty subjects of one oddball experiment, and the widget lets you walk the grid: pick a path, read its estimate and its p-value, see where that path sits in the distribution of all the others, and read the per-decision table that says which choice moves the answer.
It runs in two modes over the same six decisions, and the contrast between them is the point.
null— the condition labels have been permuted within each subject, so the expected difference is exactly zero at every channel and latency. Every path that reaches α here is a false positive, by construction. The counts at the top say how many did, what an equal number of independent tests would have given, and — from the producer’s 200-draw sweep — how often at least one of the 486 paths reached α against how often a single pre-specified path did.multiverse— the same 486 pipelines on the real contrast, where the effect is genuinely present. The headline is the median estimate, the full range and the proportion reaching α, reported together, which is how a multiverse is reported.
Nothing here implies that anyone cheated. Every option in the grid has a published precedent, which the data file states, and the point the widget exists to make is the size of the search space — the number of answers one honest analyst could defensibly report from one dataset.
Controls
| Control | What it sets |
|---|---|
| Mode | null (data with nothing to find) or multiverse (the real contrast) |
| Reference | Average of the EEG channels, linked mastoids, or Cz |
| High-pass cutoff | 0.01, 0.1 or 0.5 Hz |
| Baseline window | −200 to 0 ms, or −100 to 0 ms |
| Artifact rejection | 75 µV or 100 µV peak-to-peak, or none |
| Measurement window | 300–500, 300–600 or 350–650 ms |
| Electrode | Pz, CPz, or the Pz + CPz + Cz cluster |
| Take a different route | Re-draw the path at random, leaving the dataset alone |
| Draw another null dataset | Step to the next shipped draw (null mode) |
| Walk the pre-specified path | Jump to the pipeline the producer fixed in advance |
| α | The significance level the counts and the reference rule use |
| Outcome measure | Whether the histogram and the curve show the p-value or the estimate |
| Bins | Histogram resolution |
| Show marginals / show table | The per-decision table, and the full outcome list as numbers |
Two deliberate design decisions are worth knowing before you read the charts. α is drawn as a thin dashed reference line with a label, and no mark is coloured differently on either side of it — a p of .049 and a p of .051 are almost the same evidence, and a chart that paints them differently teaches that they are not. And the p axis is linear from 0 to 1, because a log axis would fill the plot with the region below .05 and make small p-values look enormous.
What to look for
null
- Read the two rates side by side: how often at least one of the 486 paths reached α across the sweep, and how often the single pre-specified path did. The second is the calibration check — a fixed analysis of a true null must be wrong at about α.
- The α yardstick beside the count (486 × α = 24.3) is what you would expect from that many independent tests. These paths share the data, so they are far from independent, and the observed counts swing far above and below it.
- Draw another dataset several times. Most draws produce nothing; a few produce dozens; the worst in the sweep produced 220 of 486. Forking false positives arrive in floods, not as a drizzle.
- Change one decision at a time and watch the p-value move while the data stand still. Read the rationale printed under each control: these are not straw men.
- Pick a path before you look, and keep it across draws. Nothing about the analyst changes between the two rates — only the order of looking and choosing.
multiverse
- The headline is three numbers, not one: the median estimate, the full range, and the proportion below α. Any one of them alone is a multiverse reported dishonestly.
- Find your own path on the specification curve. The question is not “is my number right” but “where does my number sit in the range this data supports”.
- Work the per-decision table. One of the six decisions moves the answer by more than the effect itself; the other five barely move it, and you cannot tell which from first principles.
- Look at the paths that do not reach α. They are the same data analysed differently, not mistakes. “Robust in sign, variable in size” is a stronger and more useful claim than a single p.
- Switch between the two modes. Same grid, same decisions, same widget — the only difference is whether the labels were permuted.
Used in
- L6.1 The multiple-comparisons landscape (
null) - L6.4 Circularity and analytic flexibility (
multiverse)
What the shipped grids show
ds-erpcore, the P3 active visual oddball, contrast target − standard mean amplitude, 20 subjects, 486 paths over six decisions (3 × 3 × 2 × 3 × 3 × 3). Held fixed so the grid stays readable: a 30 Hz low-pass, epochs −0.2 to 0.8 s at 256 Hz (from 1024 Hz), ocular correction by regression on bipolar EOG derivations, mean amplitude as the measure, a two-sided paired t-test across subjects, and a 15-trial-per-condition minimum. The file says so, and adds that the real space of defensible pipelines is larger than 486 — which is itself part of the lesson.
The null grid. Within each subject the vector of condition labels is permuted; the epochs, their order, their artifacts, their timing and the 40/160 target-to-standard counts are untouched, and only which epoch is called “target” changes. Within one draw the permutation is shared by all 486 paths, because that is the situation an analyst is in — one dataset, many analyses. Across draws the permutations are independent, which is what makes the sweep a rate rather than an anecdote.
| Quantity | Value |
|---|---|
| Paths reaching α in draw 0 | 0 of 486 (the independent-test yardstick is 24.3) |
| p across paths, draw 0 | 0.0523 … 0.9997, median 0.6332 — the nearest path misses α by 0.002 |
| Estimate across paths, draw 0 | −0.32 … +0.63 µV, median +0.10 µV — a 0.95 µV spread with nothing to find |
| Shipped draws 0–5, significant paths | 0, 8, 21, 0, 37, 1 |
| At least one path below α | about 76 % of draws (200 draws give 79.5 %, 2,000 give 75.4 %; the 200-draw 95 % interval is 74–85 %) |
| Over 200 draws: the single pre-specified path | 4.0 % of draws |
| When any path fired, how many | median 18; the worst draw, 220 of 486 |
The multiverse grid, the same 486 pipelines on the real contrast:
| Quantity | Value |
|---|---|
| Paths reaching α | 476 of 486 |
| Estimate across paths | +0.40 … +6.05 µV, median +2.32 µV; every path positive |
| p across paths | below 1e−7 … 0.2051, median 0.0004 |
| Reference marginal (median estimate) | linked mastoids +5.26 · average +2.32 · Cz +0.96 µV |
| Electrode marginal | Pz +2.67 · CPz +2.32 · Pz+CPz+Cz +2.08 µV |
| High-pass marginal | 0.01 Hz +2.33 · 0.1 Hz +2.35 · 0.5 Hz +2.27 µV |
| Window marginal | 300–500 +2.26 · 300–600 +2.34 · 350–650 +2.38 µV |
The effect is real and nearly every defensible pipeline finds it — and the reported size runs over a factor of fifteen, with one decision, the reference, accounting for more of the spread than the effect’s own median. Note that the proportion below α is not a corrected p-value and must not be read as one: the paths share the data, so 476 of 486 is one dataset agreeing with itself.
Data provenance
ds-erpcore — ERP CORE (Kappenman et al., 2021), open access, per-subject downloadable, P3 paradigm, 20 subjects. Derived assets: epoched, filtered, ocular-corrected, re-referenced, resampled, measured and tested; nothing raw is shipped. Written by data/scripts/make_forking_paths.py and registered in data/manifest.json as paths.json (an index), paths-null.json and paths-multiverse.json. The frame’s provenance line is the authority.
The grid files carry more than the outcomes: the choice set with a one-line rationale for every option, the parameters held fixed, the construction of the null in the producer’s own words, the spread, six full draws, and the 200-draw summary with the pre-specified path written out. All of it is readable from the file, which is deliberate — a multiverse whose decision set is not inspectable is not a multiverse.
TODO(confirm): ERP CORE’s licence is contested at source, and all three statements were verified from the primary material on 2026-09-18: the LICENSE file shipped with the data says CC BY-SA 4.0 with an explicit share-alike clause, dataset_description.json says CC0, and the OSF node record says CC BY 4.0. Under §10.7’s most-restrictive rule the site records CC BY-SA 4.0, and the author’s decision of 2026-09-18 (§13 item 15) permits derived assets to ship provided each carries CC BY-SA 4.0 and the attribution the LICENSE file asks for — which complies under all three readings. The author still reconciles the three statements and mirrors the entry into the catalogue registry (§10.11 item 8).
If the shipped grids are ever absent the widget falls back to a seeded synthetic demonstration — 144 paths over a simulated study, with the null true by construction — and says so on screen. A synthetic fallback is a teaching illustration, not a result, and its numbers differ from the ones above.
Open the code
site/src/components/widgets/w-garden-of-forking-paths/ — Widget.svelte, index.ts, meta.ts, state.ts, paths.ts, chart.ts, stats.ts, synthetic.ts, README.md. Repository link: TODO(confirm) (GitHub org/repo, §13 item 3).