Circularity and analytic flexibility

Double-dipping, orthogonal contrasts and collapsed localizers, preregistration, and multiverse analysis.

~60 min Widget: w-garden-of-forking-paths Notebook: nb-6-4-multiverse

Prerequisites: L6.1 · The multiple-comparisons landscape

1 claim on this page is unverified. TODO(confirm) marks a specific statement the author has not yet checked against a primary source. Everything else on this page has been reviewed. Treat a marked claim as provisional and go to the cited source rather than quoting the sentence.

Objectives

  • Recognize double-dipping (windows and channels chosen from the tested data)
  • Use orthogonal contrasts, collapsed localizers and cross-validation for selection
  • Write a preregistration
  • Run a multiverse analysis

Why this matters

L6.1 showed 486 defensible pipelines finding an effect in data that had none, in four datasets out of five. This lesson runs the same 486 pipelines on the real contrast, where the effect is genuinely there — and 476 of them reach p < .05, every single one of them positive, with the reported size running from +0.40 to +6.05 µV around a median of +2.32. A factor of fifteen between the smallest and the largest defensible answer to one question about one dataset. The reference electrode alone moves the median from +0.96 µV to +5.26 µV.

That is the honest picture of a robust effect, and it is the reason a single number with two decimal places is a smaller claim than it looks.

Concepts

Circularity: selecting and testing on the same data

A circular analysis uses the data to choose what to test, and then tests it on the same data ( (Kriegeskorte et al., 2009) ). The test’s null distribution is computed as if the choice had been made independently, so the reported p-value belongs to an analysis nobody ran.

The forms it takes in EEG are specific and worth being able to name:

  • The window from the difference wave. Plot the condition difference, find where it looks biggest, measure there. This is pf-post-hoc-windows and it is the most common form.
  • The electrode from the difference wave. Same mechanism, spatial axis.
  • Measuring inside a significant cluster. The cluster’s boundaries are a property of the observed data at a chosen threshold (L3.7); measuring the effect inside them and testing again tests the same data twice.
  • The band from the contrast. Choosing 8–13 Hz because that is where the two conditions separated in this sample.
  • Subject or trial selection by outcome. Excluding subjects who “did not show the effect” is the purest form, and it is still occasionally defended.
  • Choosing the preprocessing by the result. A reference, a filter cutoff or a rejection threshold selected because the effect came out cleaner. Each of those is a defensible option until the data pick it.

The common structure: a selection was made from a set of alternatives, using the same data that the test then treats as fresh. The severity scales with how many alternatives you would have accepted — which, as the next section says, is usually far more than the number you tried.

The non-circular ways to make the same choices

The constructive half. Every selection above has a version that does not spend the test’s error budget:

  • Pre-specification. Fix the window, electrodes, measure, band and exclusion rule from prior literature or your own earlier data, before unblinding. The cheapest option and the most powerful when the choice is right.
  • An orthogonal contrast. Select using a contrast that is independent of the one you test — under the null, it carries no information about the tested difference. The classic case is selecting on a main effect and testing an interaction, and it requires an argument for independence, not just a different-looking contrast.
  • A collapsed localizer. Choose the window and electrodes from the grand average of all conditions together. Under the null that average is uninformative about the condition difference, so the selection stays orthogonal (L3.3). Declare that you used one, because it is a choice with consequences: it finds the component, not the effect, and those can sit in different places.
  • An independent localizer. A separate block, a separate task, or a separate session, analysed on its own.
  • Cross-validation. Select on the training folds, measure on the held-out fold, and iterate. This is standard in decoding (L6.5) and works just as well for choosing a window — provided every label-aware step sits inside the fold, which is the whole content of pf-decoding-leakage.
  • Leave-one-subject-out selection. For group analyses, choose each subject’s window from the grand average of the other subjects. Simple, cheap, and rarely used.
  • Test the whole space instead of choosing. A cluster-based permutation test over the full epoch makes the search explicit and pays for it (L6.1).

Forking paths: the inflation that arrives without anyone searching

(Gelman, 2014) ‘s contribution is the observation that all of this happens even when only one analysis is ever run. The analyst does not try 486 pipelines and pick the best. The analyst looks at the data once, makes a sequence of reasonable decisions — this reference, because the drift looked bad; this rejection threshold, because a few trials were dreadful; this window, because the component peaked a little late in this sample — and reports the single result. Each decision was defensible on its own terms and was made in good faith.

The inflation arrives anyway, because the decisions were contingent on the data. Had the data looked different, different defensible decisions would have followed, and the set of analyses that could have been reported is what the error rate has to be computed over. That set is the garden, and roughly three in four against about one in twenty is its size.

Two things follow, and both matter more than the arithmetic:

  • This is not an accusation. Nothing in the widget’s 486 paths is a straw man; the data file states the published precedent for every option. A careful, honest analyst choosing one at a time is exactly who that rate is about, which is why vigilance and integrity are not the remedy. Procedure is.
  • The remedy is structural. Fix the choices before the data can influence them, or report what the whole set of choices produces. Those are preregistration and multiverse analysis, and they are the rest of this lesson.

Multiverse analysis: report the grid, not the cell

A multiverse ( (Steegen et al., 2016) ) runs every reasonable pipeline and reports the distribution of results. The steps:

  1. Enumerate the decisions and, for each, the options a reviewer would accept. Justify the set — both what is in it and what is out.
  2. Fix everything else. A multiverse is only readable if the varied decisions are the ones under discussion; the shipped grid holds the epoch, the low-pass, the ocular correction, the measure and the group test constant, and says so.
  3. Run every combination and record the estimate, the interval and the p-value for each.
  4. Report the distribution: the median estimate, the full range, the proportion reaching α, and a specification curve — every path sorted by its estimate, with your own path marked.
  5. Report the marginals: for each decision, the outcome under each of its options averaged over all the others. This is what tells you which choices matter.
  6. Mark your pre-specified path. A multiverse is a robustness report around a primary analysis, not a replacement for one. Without a declared primary path, “we ran everything” becomes another way of choosing afterwards.

The shipped grid is 486 pipelines over six decisions — three references, three high-pass cutoffs, two baselines, three rejection criteria, three a-priori windows, three electrode choices — on 20 subjects of the ERP CORE P3 paradigm, target minus standard. What it shows:

QuantityValue
Paths reaching p < .05476 of 486
Estimate across paths+0.40 to +6.05 µV, median +2.32 µV; every path positive
Reference marginal (median estimate)linked mastoids +5.26 µV · average +2.32 µV · Cz +0.96 µV
Electrode marginalPz +2.67 · CPz +2.32 · Pz+CPz+Cz +2.08 µV
High-pass marginal0.01 Hz +2.33 · 0.1 Hz +2.35 · 0.5 Hz +2.27 µV
Window marginal300–500 ms +2.26 · 300–600 ms +2.34 · 350–650 ms +2.38 µV

Read those last three rows against the first. The high-pass cutoff, the measurement window and the electrode barely move the answer; the reference moves it by more than four microvolts, which is more than the median effect itself. That is not a defect in the analysis — it is a fact about what a reference is (L2.3, and pf-reference-changes-everything): every measurement is a difference against the reference, so how much of the effect survives depends on how much of the effect the reference itself carries. The P3b is centro-parietal, so Cz carries a great deal of it and referencing to Cz subtracts most of it away; the mastoids carry comparatively little, so a mastoid reference leaves the most; the average of 30 channels sits between them. Same brains, same trials, three defensible answers. Which means the useful output of a multiverse is not “the effect is robust” but “here is which decision a reader needs to know about”, and for a P3 amplitude that decision is the reference.

Judgment call

A multiverse is a description, not a test. The proportion of paths reaching α is not a corrected p-value, and it must not be reported as one: the paths share the data, so they are strongly dependent, and 476 of 486 is one dataset agreeing with itself, not 476 replications. What the grid supports is a statement about robustness — that the sign of this effect does not depend on any of the six decisions, while its magnitude depends heavily on one of them. Say the whole sentence, or the number gets read as a test.

Preregistration

A preregistration is a time-stamped statement of the analysis, written before the data can influence it. It does not make an analysis correct, and it does not prevent exploration — it makes the distinction between confirmatory and exploratory legible, which is the thing that cannot be reconstructed afterwards.

What an EEG preregistration has to contain to be worth writing:

  1. The hypothesis, stated so it could fail — a direction and a component, not “we will investigate”.
  2. The design: conditions, trial counts, counterbalancing, and the sample size with its justification (L6.3).
  3. Exclusion rules for subjects, trials and channels, with numeric criteria, fixed in advance and applied blind to the effect.
  4. The pipeline: filter type and cutoffs, reference, artifact handling, ICA criteria, epoch and baseline — the parameters, not “standard preprocessing”.
  5. The measurement: component, electrodes, window, measure, and where the window came from.
  6. The test: the model, the unit of analysis, the correction, α, tails, and — for a permutation test — the scheme, the count and the seed.
  7. A stopping rule, if data collection is sequential.
  8. What would count as disconfirmation, and what you will do if the analysis cannot be run as written.

The last point is the one people leave out and then need. Preregistrations meet reality; the requirement is not that you never deviate, but that deviations are reported as deviations, with the reason and with the pre-registered analysis shown alongside. A registered report, where the plan is peer-reviewed and accepted before the data exist, extends the same logic to the decision to publish.

Ten lines is enough. The exercise below asks for exactly that.

The data behind this lesson

  • ds-erpcore P3, 20 subjects, target-minus-standard mean amplitude, 486 paths; epochs −0.2 to 0.8 s at 256 Hz, 30 Hz low-pass, ocular correction by regression, two-sided paired t-test across subjects. The grid file carries the choice set, the rationale for every option, the parameters held fixed and the spread. Generated by data/scripts/make_forking_paths.py.
  • The same file’s null counterpart — the one L6.1 uses — is built by permuting condition labels within subject, so the two grids differ only in whether there is anything to find.
  • TODO(confirm): ERP CORE’s licence is contested at source (CC BY-SA 4.0 in the shipped LICENSE file, CC0 in dataset_description.json, CC BY 4.0 in the OSF node record). Derived assets ship under the strictest reading per the author’s decision of 2026-09-18 (§13 item 15); the author reconciles the three.

Explore

Garden of forking paths, multiverse mode — 486 defensible pipelines on an effect that is really there

mode: multiverse Open lab page →
Loading Garden of forking paths, multiverse mode — 486 defensible pipelines on an effect that is really there…

What to look for

Data provenance is recorded in each asset's sidecar under /data/widgets/w-garden-of-forking-paths/.

What to look for:

  • The headline is three numbers, not one: the median estimate, the full range and the proportion below α. A multiverse reported as any one of them alone is a multiverse reported dishonestly.
  • Find your own path on the specification curve. Set the six controls to the pipeline you would have run, and watch the marker land somewhere inside a spread of 5.65 µV. The question a multiverse answers is not “is my number right” but “where does my number sit in the range this data supports”.
  • Work the per-choice table. Switch the reference between Cz and linked mastoids and watch the median move by more than four microvolts; then switch the high-pass or the window and watch nothing happen. One of your six decisions is load-bearing and five are not, and you cannot tell which from first principles.
  • Look at the ten paths that do not reach α. They are not mistakes; they are the same data analysed differently. An effect that survives 476 of 486 defensible pipelines is robust in sign — and that is a different, weaker and more useful claim than “p < .05”.
  • Switch to null mode and back. Same grid, same decisions, same widget; the only difference is whether the labels were permuted. The contrast between the two pictures is the lesson of this level.

Practice

Circularity and analytic flexibility: the false-positive cost of choosing a window and electrode from the tested data, two fixes measured against it, and 486 defensible pipelines on one contrast nb-6-4-multiverse

Level 6 ~7 min
notebooks/L6/nb-6-4-multiverse.ipynb

Downloads from ds-erpcore.

Open in Colab Download Read it here

A small multiverse over filter, reference and window choices for the P3 effect, built from scratch so the machinery is visible: enumerate the decisions, loop the pipeline, collect the estimates, draw the specification curve, tabulate the marginals, and mark the pre-specified path. The final cell prints the spread the exercises refer to.

Exercises

Exercise ex-6-4-multiverse-spread

Numeric

From the widget in multiverse mode: how many of the 486 paths reach p < .05, what is the median estimate across all paths, and what is the largest estimate any path produces?

paths / µV / µV
paths / µV / µV
paths / µV / µV

Accepted within ±0 / 0.05 / 0.05 paths / µV / µV.

Exercise ex-6-4-reference-marginal

Numeric

Still in multiverse mode: holding every other decision as it comes, by how many microvolts does the choice of reference alone move the median estimate between its cheapest and its most generous option?

µV

Accepted within ±0.05 µV.

Exercise ex-6-4-non-circular-selection

Multiple select

You must choose a measurement window for a P3 effect, and the component's latency in your sample is genuinely uncertain. Select every selection rule that leaves the subsequent test's error rate intact.

Options (select all that apply)

Exercise ex-6-4-preregistration

Free response

Write a ten-line preregistration for this hypothesis: 'In a visual oddball task, rare target stimuli elicit a larger positive deflection than frequent standards over centro-parietal sites between 300 and 600 ms after stimulus onset.' One line per numbered item; be specific enough that someone else could run it without asking you a question.

Pitfalls

Pitfall

Windows chosen from the tested data

Symptom
Effect only in the window that looked biggest.
Cause

A measurement window, an electrode set, a measure and a polarity are all selections. Making a selection from the same data you then test means the test no longer has the error rate it claims, because the selection has already searched over alternatives the test knows nothing about.

Detect
  • Read the methods for where the window came from. “Determined from the grand average of the condition difference” is the failure; “determined from the grand average collapsed across conditions” is not. - Count the choices you would have accepted, not the ones you made. If more than one window, electrode set or measure would have been defensible, the analysis has degrees of freedom the p-value do…
Fix
  • Fix the window in advance, from published work or from your own previous data, and record it before unblinding. - Use a collapsed localizer when the component’s latency in your sample is genuinely uncertain: choose the window and electrodes from the average of all conditions together, which under the null carries no information about the condition difference, and then measure the difference the…

Full entry with example →

Reading

  1. Kriegeskorte et al. (2009). Circular analysis. unverified
  2. Gelman & Loken (2014). Garden of forking paths. unverified
  3. Steegen et al. (2016). Multiverse analysis. unverified