Statistics for time-frequency
Cluster-permutation tests over channel × time × frequency, TFCE, limiting the search space with a priori ROIs and bands, and reporting.
Prerequisites: L3.7 · Statistics for ERPs, L4.4 · Baseline normalization and ERD/ERS
1 claim on this page is unverified. TODO(confirm) marks a specific statement the author has not yet checked against a
primary source. Everything else on this page has been reviewed. Treat a marked claim as
provisional and go to the cited source rather than quoting the sentence.
Objectives
- Run cluster-permutation tests over channel × time × frequency
- Apply TFCE
- Use a priori ROIs and bands to limit the search space
- Report TF statistics correctly
Why this matters
Nothing about cluster-based permutation testing changes when a frequency axis is added — the permutation scheme, the max-statistic null and what the p-value licenses are all exactly as L3.7 left them. What changes is the size of what you are searching and the number of decisions you made before you searched it. A 30-channel epoch at 250 Hz over one second is 7,500 tests; add 40 frequency rows and it is 300,000, and every one of the parameters that produced those rows — the estimator, the cycle scheme, the baseline window, the normalization — was a choice. The statistics are the same; the discipline required to use them honestly is greater.
This lesson is the third stop on the statistics thread. L3.7 built the cluster-permutation test and stated precisely what its p-value licenses; this lesson adds a dimension to the search space; L6.1 places both in the wider multiple-comparisons landscape (including TFCE and FDR), and L6.4 quantifies what the parameter choices above cost when they are made with the result in view.
Concepts
The same procedure, one dimension up
The recipe from L3.7, unchanged:
- A statistic at every point — now every channel × time × frequency point.
- Points above a cluster-forming threshold, grouped into connected components.
- One number per cluster, conventionally the sum of the statistic inside it.
- Per permutation, keep only the largest cluster statistic anywhere in the map.
- Compare the observed clusters against that distribution of maxima.
Step 4 is still what makes the result corrected: the null is a distribution of maxima over the entire search space, so a cluster that beats it has beaten everything the analysis looked at, however large that was. Adding a frequency axis does not weaken the family-wise error control — it is α by construction either way.
What it does change is sensitivity. A bigger space gives every permutation more opportunities to produce a large cluster, so the null shifts upward and a real effect has to be bigger, or more extended, to stand out against it. That is the honest price of searching more, and it is why the pre-specification in the next section is worth power rather than merely being tidy.
Three-dimensional adjacency
A cluster is a set of points that are adjacent, and with three axes adjacency has to be defined on all three:
- Time: neighbouring samples, implicitly.
- Frequency: neighbouring rows of the frequency grid, implicitly — but note that the rows are not independent. The estimator smooths over a bandwidth (L4.2, L4.3), so neighbouring rows of a Morlet map at 7 cycles share most of their data, and a logarithmic frequency grid has a different amount of overlap at each end. The adjacency graph treats them as neighbours; the smoothing already made them nearly the same measurement.
- Channels: an explicit neighbourhood graph built from the sensor positions — a triangulation of the montage, or a distance criterion. It is a modelling choice, it has to be stated, and it belongs in the methods section with the montage that produced it.
Two consequences follow from the frequency axis in particular. First, an effect that is genuinely narrowband can be spread across many rows by the estimator’s own smoothing and appear as a large cluster, which flatters its extent. Second, choosing a finer frequency grid does not add information — it adds correlated rows — so it inflates the point count of every cluster, in the observed map and in every permutation alike. The test remains valid; the description of the cluster becomes even less like a measurement than it was in two dimensions.
Threshold-free cluster enhancement
The cluster-forming threshold is an analysis choice that changes cluster extent without changing the error rate (L3.7), which makes it uncomfortable: a high threshold favours focal, strong effects and a low one favours weak, extended effects, and sliding it until something appears is a search. TFCE removes the choice by scoring each point according to the support it receives from clusters at every threshold, integrating the cluster extent against threshold height with fixed exponents, and then running the same max-statistic permutation over the resulting image. (Smith, 2009) is the original treatment; L6.1 places it beside the other corrections.
Two things TFCE does not do. It does not remove the other parameters — the exponents and the step size are still choices, usually left at their published defaults and therefore worth stating. And it does not change what the result licenses: a significant TFCE point still belongs to a family-wise-corrected test of the whole space, and reading the surviving region as an onset, an extent or a location is the same error it was before (pf-cluster-inference-misread).
Limiting the search space in advance
Because the search space is now large enough to cost real power, restricting it in advance is worth more here than anywhere else in the curriculum — and it is only legitimate when the restriction genuinely precedes the data:
- A region of interest. C3 and C4 for a hand-movement ERD; a posterior cluster for alpha. Justified from the literature or from an independent localizer, fixed in the analysis plan, and reported whether or not it worked.
- A band. Mu, beta, or an individualized band defined by each subject’s own peak (L4.6) — which is itself a form of pre-specification, since the band is set by the subject’s spectrum rather than by the effect.
- A time window. Taken from the paradigm, not from the map.
- A single summary per subject. The extreme case: average within the ROI, the band and the window, and run one paired test on the resulting numbers. No cluster test, no correction, full power for exactly one hypothesis, and blindness everywhere else.
The multiplicity that is easiest to forget is the time-frequency parameters themselves. Estimator, cycle scheme, frequency grid, baseline window, normalization, averaging order (L4.4) — each is a choice, each changes the map, and trying several and reporting the one that produced a cluster is pf-post-hoc-windows with more dimensions. Fix them in the analysis plan; if you must explore, use a collapsed localizer or a separate dataset, and report the exploration.
Write down, before you compute the maps: the estimator and its parameters as formulas, the baseline window and the normalization, the channels, the band, the time window, the cluster-forming threshold or TFCE settings, the permutation scheme and count, α, the tails, and the one sentence you intend to publish. Then check that sentence against L3.7’s list of what a cluster p-value licenses. If it contains a frequency — “the effect was in the beta band” — that is now a location claim on the new axis, and the test does not support it any more than it supports an onset.
Reporting
Everything L3.7 asked for, plus the time-frequency layer:
- the estimator, its parameters per frequency, and the frequency grid;
- the epoch length actually transformed, the edge margin and whether it was excluded from the tested space (L4.2 — testing inside the edge margin tests padding);
- the baseline window, the normalization, and the order of averaging;
- the tested space: which channels, which frequencies, which times;
- the adjacency definition on all three axes, including how channel neighbours were derived;
- the point-wise statistic, the cluster-forming threshold or the TFCE settings, the permutation scheme, the number of permutations and the seed;
- every cluster with its statistic and p-value, including the ones that did not survive, and the extent described as an extent;
- the smallest attainable p-value, which is 1/(number of permutations) because the identity relabelling counts as one of them — so 1,024 permutations put the floor at .001, and a p there means “below this null’s resolution”, not “essentially zero” (the same floor L3.7 reads off the shipped fixture).
The data behind this lesson
- The notebook runs a cluster test on left- versus right-hand imagery ERD maps from
ds-eegbci(ODC-By 1.0, open access; 64 channels, 160 Hz, no hardware filters), with the estimator, the baseline, the tested space, the adjacency, the permutation count and the seed all stated in one cell. - The widget’s
tfmode runs a real three-dimensional test. 64 channels × 15 frequencies (6–34 Hz) × 23 times (−0.5 to 3.9 s) is 22,080 point-wise tests on left- minus right-hand imagery fromds-eegbci, 20 subjects, cluster-forming threshold 2.0930 (the two-tailed .05 point of t on 19 df), 1,024 sign-flip permutations with a recorded seed, and an adjacency built bycombine_adjacencyover all three axes from a Delaunay triangulation of the montage. The exact call is in the widget’s sidecar. - The same data is tested twice, and the second test is the one to look at. Over the whole space there are 341 candidate clusters and one survives, Σt = +591.2 over 200 points and 35 channels, 6–30 Hz, 0.1–1.7 s, p = .027. Restricted in advance to a sensorimotor ROI and the 8–13 Hz band — 1,449 tests, a family 15.2× smaller — there are 31 candidates and none survives, the best at p = .160. That is the honest result and it is not the tidy one. Pre-specification bought a smaller search space; it did not buy a significant effect, and it cannot, because the effect either is or is not there. Quoting either number without the other teaches the opposite of this lesson: the first alone says “cluster tests find things”, the second alone says “pre-specification kills your result”. Together they say what is true — the correction is doing its job in both cases, and a smaller family is a claim you can defend, not a claim you are more likely to make.
- For scale: 1,511 of the 22,080 points exceed the threshold uncorrected, where chance alone predicts about 1,104.
- TFCE is not shipped. It needs its own null distribution and did not fit the budget for this phase, so the widget names the option and does not relabel these numbers as TFCE.
TODO(confirm): an author review should decide whether Level 4 wants a TFCE fixture of its own.
Explore
Work the tf panel first. Turn the uncorrected map on and count the points above threshold against the number chance predicts, then notice that the count is close to it — 1,511 against about 1,104 — which is what “nothing survives correction” looks like before correction. Step the permutation player and watch each relabelled dataset contribute exactly one maximum to the null, however many points the map holds; that single fact is why adding a frequency axis does not cost you power the way a Bonferroni correction would. Move the cluster-forming threshold and watch the clusters change shape while the p-value does not collapse, because the null moves with them.
Then switch the tested space to the pre-declared ROI and band and watch what happens: the family shrinks by a factor of 15, the null narrows with it, and the surviving cluster does not reappear. Sit with that. Pre-specification is a claim about what you were entitled to look at, not a technique for getting a smaller p-value.
The erp panel runs the identical procedure on a two-dimensional map, where the whole tested space fits on one screen. Use it when the three-dimensional view is hard to read, and read the “what this licenses” panel in both.
Practice
Statistics for time-frequency: uncorrected cell counts, cluster-permutation and TFCE over frequency and time, and three analysis plans ranked by analyst degrees of freedom nb-4-7-tf-cluster
Downloads from ds-eegbci.
The notebook runs a cluster-permutation test over channel × time × frequency on ds-eegbci motor-imagery ERD maps, with the full parameter list and the seed in one cell, and prints the cluster table including the clusters that did not survive.
Exercises
Exercise ex-4-7-rank-plans
Put in orderThree analysis plans for the same left-versus-right imagery contrast. Put them in order from the MOST degrees of freedom left to the analyst to the FEWEST.
- Fix the estimator, cycle scheme, baseline window and normalization in advance; compute maps over all 64 channels, 1–100 Hz and the whole epoch; run one cluster-permutation test over that whole space with a cluster-forming threshold fixed in advance.
- Compute time-frequency maps over all 64 channels, 1–100 Hz and the whole epoch; try cycle counts of 3, 5 and 7, three baseline windows and both dB and percent normalization; run a cluster test on each combination and report the one that yields the smallest cluster p-value.
- Pre-register the channel ROI (C3 and C4), the band (8–13 Hz), the time window (0.5–3.5 s), the estimator and cycle scheme, the baseline window and the normalization; average within the ROI, band and window to one number per subject per condition; run one paired t-test on those numbers.
Exercise ex-4-7-frequency-axis
Multiple choiceYou extend a cluster-permutation test from 30 channels × 129 time points to the same channels and times × 40 frequency rows, changing nothing else. What happens to the test?
Exercise ex-4-7-rewrite-tf-claim
Free responseA draft reads: 'A cluster-based permutation test showed a significant beta-band desynchronization over left sensorimotor cortex beginning 300 ms after the cue (p = .001).' Rewrite it to say only what the test supports, and list what else the methods section needs.
Pitfalls
Mass univariate tests without correction
- Symptom
- Significant blobs everywhere.
- Cause
Testing every channel and every time point at α = .05 means running one test per point and accepting a five per cent false-positive rate on each. A 30-channel epoch sampled at 250 Hz over one second is 7,500 tests, so about 375 points are expected to pass by chance when nothing at all is happening. That is not a defect of the data; it is the definition of the α level applied 7,500 times.
- Detect
- Count. Multiply the number of channels by the number of time points, multiply by α, and compare with the number of significant points you found. If they are similar, you have measured your α level. - Look at the baseline. Significant points before the stimulus, where the conditions cannot yet differ, are chance points by construction. - Run the analysis under the null. Sign-flip the subject-lev…
- Fix
- Reduce the search space before testing. A single pre-specified measurement in a fixed window at fixed electrodes needs one test and no correction (L3.3, L3.7). - Use a cluster-based permutation test when the window cannot be pre-specified. Because the null is a distribution of maxima over the entire searched space, one p-value covers the whole search and the family-wise error rate is controlled…
Cluster p-values read as latency/location claims
- Symptom
- "The effect began at 312 ms" from a cluster test.
- Cause
A cluster-based permutation test performs one test of one hypothesis: that the condition labels are exchangeable across the entire tested set of channels and time points. The null distribution is built from the largest cluster statistic found anywhere in each permuted map, which is exactly what makes the single p-value valid over the whole search — and exactly what prevents it from saying anythin…
- Detect
- Read the claim back against the test. Any sentence containing an onset, an offset, a duration, a location or a comparison between clusters is not supported by a cluster p-value. - Move the cluster-forming threshold and watch the boundaries. If the reported onset changes by tens of milliseconds across thresholds you would all have accepted, it is a property of the threshold. - Check whether the…
- Fix
- Report the cluster as a description, in those words: “the cluster extended over these channels between these latencies, with its largest statistic at …”, followed by the single inferential claim the test supports — that the conditions differ somewhere in the tested space, with the family-wise error rate controlled at α. - Report the full analysis: the tested space, the point-wise statistic, the…
In other tools
In other toolsEEGLAB · FieldTrip — names only
The equivalents of what this lesson does, for a reader who works in another toolbox. Function names only: their own documentation is the place to learn how to call them.
EEGLAB
statcondEEGLABstd_erspplotEEGLAB
FieldTrip
ft_freqstatistics(montecarlo, cluster)FieldTripft_clusterplotFieldTrip
Names checked 2026-09-18 against EEGLAB 2026.0.0 (plugins at the versions in EEGLAB’s own plugin list) and FieldTrip 20251218.
Reading
- Maris & Oostenveld (2007). Nonparametric testing. unverified
- Smith & Nichols (2009). TFCE. unverified