P2 precomputed w-garden-of-forking-paths

Garden of Forking Paths

Analytic flexibility under the null and on real data — 486 defensible pipelines over one oddball experiment, counted two ways.

2 claims on this page are unverified. TODO(confirm) marks a specific statement the author has not yet checked against a primary source. Everything else on this page has been reviewed. Treat a marked claim as provisional and go to the cited source rather than quoting the sentence.

Modes: null · multiverse — the widget below runs in null. Use Share state to put the exact view in the URL.

Garden of Forking Paths

mode: null
Loading Garden of Forking Paths…

Data provenance is recorded in each asset's sidecar under /data/widgets/w-garden-of-forking-paths/.

What it does

One dataset, 486 analyses, and the question of what a single reported number is worth.

The widget holds a grid of pipelines a reviewer would accept without comment — six decisions with two or three options each: the reference, the high-pass cutoff, the baseline interval, the artifact-rejection criterion, the measurement window and the electrode. Every combination has been run over twenty subjects of one oddball experiment, and the widget lets you walk the grid: pick a path, read its estimate and its p-value, see where that path sits in the distribution of all the others, and read the per-decision table that says which choice moves the answer.

It runs in two modes over the same six decisions, and the contrast between them is the point.

  • null — the condition labels have been permuted within each subject, so the expected difference is exactly zero at every channel and latency. Every path that reaches α here is a false positive, by construction. The counts at the top say how many did, what an equal number of independent tests would have given, and — from the producer’s 200-draw sweep — how often at least one of the 486 paths reached α against how often a single pre-specified path did.
  • multiverse — the same 486 pipelines on the real contrast, where the effect is genuinely present. The headline is the median estimate, the full range and the proportion reaching α, reported together, which is how a multiverse is reported.

Nothing here implies that anyone cheated. Every option in the grid has a published precedent, which the data file states, and the point the widget exists to make is the size of the search space — the number of answers one honest analyst could defensibly report from one dataset.

Controls

ControlWhat it sets
Modenull (data with nothing to find) or multiverse (the real contrast)
ReferenceAverage of the EEG channels, linked mastoids, or Cz
High-pass cutoff0.01, 0.1 or 0.5 Hz
Baseline window−200 to 0 ms, or −100 to 0 ms
Artifact rejection75 µV or 100 µV peak-to-peak, or none
Measurement window300–500, 300–600 or 350–650 ms
ElectrodePz, CPz, or the Pz + CPz + Cz cluster
Take a different routeRe-draw the path at random, leaving the dataset alone
Draw another null datasetStep to the next shipped draw (null mode)
Walk the pre-specified pathJump to the pipeline the producer fixed in advance
αThe significance level the counts and the reference rule use
Outcome measureWhether the histogram and the curve show the p-value or the estimate
BinsHistogram resolution
Show marginals / show tableThe per-decision table, and the full outcome list as numbers

Two deliberate design decisions are worth knowing before you read the charts. α is drawn as a thin dashed reference line with a label, and no mark is coloured differently on either side of it — a p of .049 and a p of .051 are almost the same evidence, and a chart that paints them differently teaches that they are not. And the p axis is linear from 0 to 1, because a log axis would fill the plot with the region below .05 and make small p-values look enormous.

What to look for

null

  • Read the two rates side by side: how often at least one of the 486 paths reached α across the sweep, and how often the single pre-specified path did. The second is the calibration check — a fixed analysis of a true null must be wrong at about α.
  • The α yardstick beside the count (486 × α = 24.3) is what you would expect from that many independent tests. These paths share the data, so they are far from independent, and the observed counts swing far above and below it.
  • Draw another dataset several times. Most draws produce nothing; a few produce dozens; the worst in the sweep produced 220 of 486. Forking false positives arrive in floods, not as a drizzle.
  • Change one decision at a time and watch the p-value move while the data stand still. Read the rationale printed under each control: these are not straw men.
  • Pick a path before you look, and keep it across draws. Nothing about the analyst changes between the two rates — only the order of looking and choosing.

multiverse

  • The headline is three numbers, not one: the median estimate, the full range, and the proportion below α. Any one of them alone is a multiverse reported dishonestly.
  • Find your own path on the specification curve. The question is not “is my number right” but “where does my number sit in the range this data supports”.
  • Work the per-decision table. One of the six decisions moves the answer by more than the effect itself; the other five barely move it, and you cannot tell which from first principles.
  • Look at the paths that do not reach α. They are the same data analysed differently, not mistakes. “Robust in sign, variable in size” is a stronger and more useful claim than a single p.
  • Switch between the two modes. Same grid, same decisions, same widget — the only difference is whether the labels were permuted.

Used in

What the shipped grids show

ds-erpcore, the P3 active visual oddball, contrast target − standard mean amplitude, 20 subjects, 486 paths over six decisions (3 × 3 × 2 × 3 × 3 × 3). Held fixed so the grid stays readable: a 30 Hz low-pass, epochs −0.2 to 0.8 s at 256 Hz (from 1024 Hz), ocular correction by regression on bipolar EOG derivations, mean amplitude as the measure, a two-sided paired t-test across subjects, and a 15-trial-per-condition minimum. The file says so, and adds that the real space of defensible pipelines is larger than 486 — which is itself part of the lesson.

The null grid. Within each subject the vector of condition labels is permuted; the epochs, their order, their artifacts, their timing and the 40/160 target-to-standard counts are untouched, and only which epoch is called “target” changes. Within one draw the permutation is shared by all 486 paths, because that is the situation an analyst is in — one dataset, many analyses. Across draws the permutations are independent, which is what makes the sweep a rate rather than an anecdote.

QuantityValue
Paths reaching α in draw 00 of 486 (the independent-test yardstick is 24.3)
p across paths, draw 00.0523 … 0.9997, median 0.6332 — the nearest path misses α by 0.002
Estimate across paths, draw 0−0.32 … +0.63 µV, median +0.10 µV — a 0.95 µV spread with nothing to find
Shipped draws 0–5, significant paths0, 8, 21, 0, 37, 1
At least one path below αabout 76 % of draws (200 draws give 79.5 %, 2,000 give 75.4 %; the 200-draw 95 % interval is 74–85 %)
Over 200 draws: the single pre-specified path4.0 % of draws
When any path fired, how manymedian 18; the worst draw, 220 of 486

The multiverse grid, the same 486 pipelines on the real contrast:

QuantityValue
Paths reaching α476 of 486
Estimate across paths+0.40 … +6.05 µV, median +2.32 µV; every path positive
p across pathsbelow 1e−7 … 0.2051, median 0.0004
Reference marginal (median estimate)linked mastoids +5.26 · average +2.32 · Cz +0.96 µV
Electrode marginalPz +2.67 · CPz +2.32 · Pz+CPz+Cz +2.08 µV
High-pass marginal0.01 Hz +2.33 · 0.1 Hz +2.35 · 0.5 Hz +2.27 µV
Window marginal300–500 +2.26 · 300–600 +2.34 · 350–650 +2.38 µV

The effect is real and nearly every defensible pipeline finds it — and the reported size runs over a factor of fifteen, with one decision, the reference, accounting for more of the spread than the effect’s own median. Note that the proportion below α is not a corrected p-value and must not be read as one: the paths share the data, so 476 of 486 is one dataset agreeing with itself.

Data provenance

ds-erpcore — ERP CORE (Kappenman et al., 2021), open access, per-subject downloadable, P3 paradigm, 20 subjects. Derived assets: epoched, filtered, ocular-corrected, re-referenced, resampled, measured and tested; nothing raw is shipped. Written by data/scripts/make_forking_paths.py and registered in data/manifest.json as paths.json (an index), paths-null.json and paths-multiverse.json. The frame’s provenance line is the authority.

The grid files carry more than the outcomes: the choice set with a one-line rationale for every option, the parameters held fixed, the construction of the null in the producer’s own words, the spread, six full draws, and the 200-draw summary with the pre-specified path written out. All of it is readable from the file, which is deliberate — a multiverse whose decision set is not inspectable is not a multiverse.

TODO(confirm): ERP CORE’s licence is contested at source, and all three statements were verified from the primary material on 2026-09-18: the LICENSE file shipped with the data says CC BY-SA 4.0 with an explicit share-alike clause, dataset_description.json says CC0, and the OSF node record says CC BY 4.0. Under §10.7’s most-restrictive rule the site records CC BY-SA 4.0, and the author’s decision of 2026-09-18 (§13 item 15) permits derived assets to ship provided each carries CC BY-SA 4.0 and the attribution the LICENSE file asks for — which complies under all three readings. The author still reconciles the three statements and mirrors the entry into the catalogue registry (§10.11 item 8).

If the shipped grids are ever absent the widget falls back to a seeded synthetic demonstration — 144 paths over a simulated study, with the null true by construction — and says so on screen. A synthetic fallback is a teaching illustration, not a result, and its numbers differ from the ones above.

Open the code

site/src/components/widgets/w-garden-of-forking-paths/Widget.svelte, index.ts, meta.ts, state.ts, paths.ts, chart.ts, stats.ts, synthetic.ts, README.md. Repository link: TODO(confirm) (GitHub org/repo, §13 item 3).