Level 6

Inference and Rigor

Make claims that survive scrutiny: control the search space, use models that respect the data structure, quantify effects and precision, and report so others can reproduce.

You can now…

  • Choose a multiple-comparisons strategy
  • Fit mixed models on trial-level data
  • Simulate power
  • Recognize and avoid circularity
  • Run and interpret decoding as inference
  • Write a methods section to COBIDAS-MEEG standard

Level 6 is about the distance between a result and a claim. Six lessons: the multiple-comparisons landscape and what each correction actually controls; mixed models, for the designs whose structure a per-subject average cannot represent; effect sizes, precision and simulation-based power; circularity and analytic flexibility, with a multiverse over 486 defensible pipelines; decoding read as inference rather than as a score; and reporting, which is the only channel through which any of it reaches a reader.

Two numbers set the level’s terms, and both come from the same 486 pipelines run over twenty subjects of one oddball experiment. On data whose null is true by construction, at least one pipeline reaches p below .05 in roughly three simulated datasets out of four, while a single pipeline fixed in advance reaches it in about one in twenty — which is what α promises, and is the check that the construction is sound. On the real contrast the effect is found by 476 of the 486, every one of them positive, with the reported size running from +0.40 to +6.05 µV. Nothing separates those pictures except when the choices were made, and nobody in any of them did anything a reviewer would call fraud. That is the point: the remedy is procedure — pre-specification, orthogonal selection, cross-validation, preregistration, reporting the whole grid — not vigilance.

This level closes the statistics thread that runs from L3.7 (testing an ERP, and what a cluster p-value licenses) through L4.7 (the same machinery over a time-frequency space) into L6.1 and L6.4. Work the lessons in order — L6.6 assumes all five before it — and finish with the capstone, which audits a published paper and then preregisters one analysis before the data are touched.

Lessons

  1. L6.1 The multiple-comparisons landscape

    Enumerate the search space; compare FWER and FDR control, Bonferroni, max-statistic, cluster permutation and TFCE; pre-specify ROIs and windows.

    ~60 min ◐ widget ▤ notebook
  2. L6.2 Mixed models and trial-level data

    What averaging-then-testing discards, linear mixed models with subject and item random effects on single trials, convergence, and reporting.

    ~75 min ▤ notebook
  3. L6.3 Effect sizes, power and precision

    Effect sizes for within-subject EEG designs, SME as precision, and simulation-based power analysis using pilot data.

    ~50 min ▤ notebook
  4. L6.4 Circularity and analytic flexibility

    Double-dipping, orthogonal contrasts and collapsed localizers, preregistration, and multiverse analysis.

    ~60 min ◐ widget ▤ notebook
  5. L6.5 Decoding as inference

    Time-resolved decoding with proper cross-validation, chance level and permutation significance, temporal generalization, and what above-chance decoding means.

    ~60 min ▤ notebook
  6. L6.6 Reporting standards

    Reporting filters, reference, rejection, ICA, epoching, measurement and statistics to COBIDAS-MEEG standard; figures with uncertainty; sharing code and data.

    ~40 min

Capstone

C6 Capstone — Rigor audit

~300 min3 deliverables