Resting-state and biomarkers
Resting-state features, test–retest reliability, and the biomarker replication problem, with the confounds of real clinical archives.
Prerequisites: L4.6 · Is it really an oscillation?, L6.3 · Effect sizes, power and precision
2 claims on this page are unverified. TODO(confirm) marks a specific statement the author has not yet checked against a
primary source. Everything else on this page has been reviewed. Treat a marked claim as
provisional and go to the cited source rather than quoting the sentence.
Objectives
- Compute resting-state features (IAF, aperiodic slope, band ratios, microstates overview)
- Estimate test–retest reliability
- Explain the biomarker replication problem
Why this matters
Resting EEG is the cheapest recording in the field and the one most often proposed as a biomarker: three minutes, no task, no training, and a number at the end. The reason so few of those numbers survive contact with a second cohort is not that the effects are absent — it is that the measurement is less reliable than the effect is large, and that the archives where the effects are found are confounded in ways no sample size repairs. This lesson computes the features, measures how reproducible they are, and reads the archives’ own documentation as evidence.
Concepts
The protocol is a variable, and it is rarely reported as one
“Resting state” names an absence of task, not a standard procedure. Compare four resting recordings this site’s directory holds:
| Dataset | Resting protocol | Sessions |
|---|---|---|
ds-lemon | 16 min: 16 alternating 1-min blocks, 8 eyes-closed and 8 eyes-open | one |
ds-srm | ~4 min eyes-closed (task resteyesc) | two, 2–3 months apart, for ~42 of 111 |
ds-dortmund | 3-min eyes-closed and 3-min eyes-open, each before and after a ~2-hour cognitive battery | two, ~5 years apart, for 208 of 608 |
ds-iowapd | ~3 min eyes-open | one |
Eyes-open and eyes-closed are different states, not two ways of doing the same thing; one minute is a different measurement from ten; and a block recorded after two hours of cognitive testing is not a block recorded at the start. ds-dortmund makes the last point unavoidable: its “pre” and “post” labels are within-session markers around that battery, not its two longitudinal sessions, and reading them as sessions is a documented way to get this dataset wrong. Report the protocol — duration, condition, order, instruction, time of day, what happened immediately before — because a feature compared across studies with different protocols is not the same feature.
Features, and what each of them actually measures
- Individual alpha frequency. The frequency of a participant’s own alpha peak, which is the anchor for individualised bands (L4.6). Two estimators are in common use — the centre of gravity of power over a fixed window, and the centre frequency of a fitted peak — and they do not return the same number on the same data. Whichever you use, a participant with no fitted alpha peak has no IAF, and “excluded” and “imputed with the group mean” give different results downstream.
- Aperiodic offset and exponent. The broadband component of the spectrum, parameterised rather than assumed away (L1.7, (Donoghue et al., 2020) ). Report the fit range, the settings and the fit quality; the range is a choice, and on a device with a hardware ceiling it is a constrained one (L7.5).
- Band power and band ratios. The most-reported resting features and the least interpretable.
pf-band-power-slopestates the mechanism: a band-power difference can be a rhythm changing, an aperiodic offset moving every band together, or an exponent change pivoting two bands in opposite directions — and a ratio is a two-point estimate of the slope with rhythms mixed in, so it is more sensitive to the aperiodic component than either band alone. If you report a band, fit the spectrum first and say whether a peak was found in it. - Microstates. Brief quasi-stable scalp topographies, clustered from the maps at peaks of global field power, then used to label the whole recording; the parameters reported are each class’s mean duration, coverage, occurrence rate and the transition probabilities between classes. TODO(confirm): the number of template classes conventionally retained, and the canonical labelling of them, are stated from general knowledge rather than from a source — §15 carries no microstates reference, so the author should add one and check both before this section leaves draft. What is safe to teach here is the structural point: microstate parameters are properties of a clustering, so the number of classes, the clustering algorithm, whether the templates were fitted per subject or per group, and the polarity convention are all analysis choices that must be reported before any parameter is comparable across papers.
Reliability is the ceiling on everything else
The intraclass correlation coefficient asks what fraction of the total variance in a measurement is between participants rather than within them. It is the quantity that decides whether a feature can be a marker of anything, and it comes in two flavours that answer different questions:
- Split-half or within-session reliability — split one recording in two and compare. This measures the estimator’s noise: how much of the number is the amount of data you happened to analyse.
ds-lemoncan supply this and cannot supply anything else, because it has one EEG session. - Test–retest reliability — the same participant on two occasions. This measures the estimator’s noise plus everything about the person and the session that changed in between, which is what a trait marker has to survive. The interval is part of the result:
ds-srm’s second session is 2–3 months later for about 42 of its 111 participants,ds-dortmund’s is about five years later for 208 of 608, and its within-session pre/post blocks give a short-interval check on the same people. A reliability figure without its interval is not interpretable.
Three consequences worth stating plainly:
- Unreliable measurement caps the observable effect. Measurement error inflates the denominator of every standardised effect size (L6.3), so a feature with poor reliability produces small effects even where the underlying difference is large — and those effects do not replicate, because their size depends on a noise term that differs between labs.
- Missing sessions are not missing at random.
ds-srmhas a second session for a subset;ds-dortmundfor a third of its cohort. Who came back is a selection, so a reliability estimate from the returners is a reliability estimate for people who return, and the count that went into it belongs in the report. - Reliability is a property of the measurement, not of the feature. The same feature, estimated from 30 seconds or from 4 minutes, on 4 electrodes or 64, with a fixed band or an individualised one, has different reliability. Report the pipeline with the ICC.
The archives, and what makes each of them hard
These are the datasets a biomarker study reaches for, and every one of them carries a documented confound. None of this is a criticism of the people who collected them — a public clinical archive is a gift, and the caveats are documented because the depositors documented them.
| Archive | What the directory records | What it means for a biomarker claim |
|---|---|---|
ds-tdbrain | 47 healthy controls among 1,274 participants; comorbidity common; two cap systems over about twenty years | The control group is 4 % of the sample, so any case–control contrast rests on 47 people; and recording era is confounded with everything that changed over twenty years, including referral patterns and diagnostic practice (pf-site-device-confound) |
ds-iowapd | 100 patients against 49 controls; patients recorded on dopaminergic medication, which attenuates beta signatures; groups differ in age and sex; Pz is the online reference and is flat | Medication is a treatment effect inside the patient group, so “patients differ from controls” and “medicated brains differ from unmedicated ones” are the same contrast (pf-group-demographic-confound) |
ds-brainlat | Mains frequency differs by country within one dataset — 50 Hz in Argentina, Chile and Peru; 60 Hz in Mexico and Colombia | Line noise, the notch that removes it, and anything measured near it differ by site, and site is confounded with country and therefore with cohort |
ds-aszed | Two amplifiers at 200 and 256 Hz with each device’s default filters, at two sites | The hardware fingerprint — bandwidth, noise floor, filter roll-off — is in the data before any analysis, and it separates the devices whether or not it separates the groups |
ds-hbn | Transdiagnostic despite the name; 3,155 participants across releases R1–R11, each a separate accession; channels are E1–E128 + Cz, not 10-20 names | It is not a healthy control cohort. Used as one, it would put a clinically heterogeneous sample on the control side of a case–control contrast |
The single most useful habit in this lesson: compute the hardware-sensitive measures first and test whether they separate the setups, before testing whether anything separates the groups. Spectral slope over the full band, power above the physiological range, line-noise residue and the noise floor are all shaped by the amplifier and the cap. If those measures separate two recording set-ups, then any group difference that follows the same split has at least two explanations, and the design decides which — not the statistics. Where a dataset has no within-setup contrast at all, as when every patient was recorded on one amplifier and every control on the other, the group effect is not estimable, and the honest report says so instead of adjusting for site and moving on. Covariate adjustment is a modelling assumption, not a repair.
Why biomarkers fail to replicate, in the order the failures happen
Put the pieces in sequence and the replication problem stops being mysterious:
- The protocol differs between the discovery cohort and the replication cohort, so the feature is not quite the same feature.
- The measurement is less reliable than anyone checked, so the effect size in the discovery sample is partly noise that will not recur.
- The discovery sample is confounded — by site, device, era, medication, age, sex — so part of the effect is real and about something other than the grouping variable.
- The analysis had flexibility that was not pre-specified: which band, which electrodes, which ratio, which covariates (L6.4).
- The result is reported as a p-value on a group difference rather than as a per-participant classification with an interval, so its distance from clinical usefulness is never visible.
The remedies are the ones this ladder has already taught, applied in the same order: report the protocol; measure reliability before effect size; tabulate hardware and demographics against group; pre-specify; and report the effect in the units of the measurement with an interval. A biomarker claim additionally needs something none of these archives can supply — prospective validation in a sample drawn the way the clinical use would draw it.
Explore
No widget is introduced here; two from Level 1 are the right instruments, and the point is to use them on a reliability question rather than a descriptive one.
- Open the aperiodic explorer in
comparemode and put two subjects side by side. Note how much of the difference between their spectra is offset, how much is exponent, and how much is the alpha peak — then ask which of the three a band-power number would have reported. - In the same widget, change the fit range and watch the exponent move. That sensitivity is a source of between-study variance that no reliability estimate computed within one pipeline will ever show you.
- Open the Welch explorer and halve the segment length. The spectrum gets noisier and the estimated peak frequency moves. That movement is the split-half reliability of IAF, seen directly, before any statistics.
- On paper: write down the exact recipe for one resting feature you would propose as a marker — condition, duration, electrodes, reference, band or fit range, estimator, exclusions. Then ask what a second lab would have to be told to produce the same number, and whether any paper you have read told you that much.
Practice
How reliable is a resting-state feature? Split-half, minutes, two months and five years of test-retest ICC for the alpha peak, a theta/beta ratio and the aperiodic slope, against a clinical group contrast and the attenuation ceiling it implies nb-7-6-reliability
Downloads from ds-lemon, ds-dortmund, ds-srm, ds-iowapd.
Split-half reliability on ds-lemon (one session, so within-session only); test–retest on ds-srm at a 2–3 month interval and on ds-dortmund at a five-year interval, with its within-session pre/post blocks as a short-interval check; then a group contrast on ds-iowapd — 100 patients against 49 controls — to ask which features are fit to be biomarkers at all. ds-tdbrain, ds-brainlat and ds-hbn are read and linked, never downloaded: the first two are behind a DUA and a registration, and ds-hbn is share-alike with no licence decision recorded for it on this site, so nothing on this site derives from it.
Exercises
Exercise ex-7-6-tdbrain-controls
NumericA proposal will use the largest psychiatric EEG archive in this site's directory for a case–control study, on the grounds that it has over a thousand participants. Read the entry for `ds-tdbrain`: how many healthy controls does it contain?
Exercise ex-7-6-archive-confounds
Multiple selectWhich of these statements are recorded in this site's dataset directory as properties of the archive they name?
Exercise ex-7-6-icc-pair
NumericFrom the notebook, for the test–retest dataset and interval it states: what is the ICC of the individual alpha frequency?
Exercise ex-7-6-fit-for-biomarker
Multiple choiceFeature A has high test–retest reliability and a small group difference. Feature B has poor reliability and a large group difference in the same discovery cohort. Which is the better candidate biomarker, and why?
Pitfalls
Band power changes that are slope changes
- Symptom
- "More beta" with no beta peak; all bands shift together.
- Cause
Band power is the area under the power spectrum in a frequency range, and the spectrum is a sum of two things: an aperiodic component, broadband and falling with frequency, described by an offset and an exponent; and whatever periodic peaks sit on top of it (L1.7). Integrating over a band adds both together and reports one number, so three physically different events are indistinguishable in it:
- Detect
- Fit the spectrum, per subject and per condition, and look at the peak list. If no peak has a centre frequency inside the band, do not name the band-power result after a rhythm. Report the fit range, the settings (peakwidthlimits, maxnpeaks, minpeakheight, peakthreshold), the aperiodic mode and the fit quality alongside. - Check whether the “peak” is at the detection floor. A peak whose fitted p…
- Fix
- Parameterize the spectrum and report the parameters — aperiodic offset and exponent (and knee where used), and every peak’s centre frequency, power and bandwidth, with the settings and the fit quality. Prefer the peak’s own parameters to band power wherever the question allows: “the alpha peak fell by 0.6 µV² and moved 0.4 Hz lower” is a claim about a rhythm; “alpha power fell” is not. - If the…
Site, amplifier or cap confounded with group or time
- Symptom
- Two amplifiers at two sites with different sampling rates; two cap systems over 20 years; 50 vs 60 Hz mains by country.
- Cause
Every acquisition choice leaves a fingerprint in the data (L0.3, L1.3): the hardware bandwidth and filters shape the spectrum, the amplifier’s noise floor sets the high-frequency baseline, the cap’s electrode set and reference change amplitudes and topographies, the sampling rate decides which harmonics are visible and how resampling behaves, and the mains frequency decides where the line and its…
- Detect
- Tabulate hardware per subject (site, amplifier, cap, rate, filters, reference, mains) from the dataset description and the file headers, and cross-tabulate it against the group and against the recording date; any cell imbalance is a confound. - Compute the hardware-sensitive measures (spectral slope over the full band, power above the physiological range, line-noise residue, noise floor) and te…
- Fix
- Design first: balance groups across sites and devices, or record every group on every setup; record era-matched controls in archives. - Harmonize what can be harmonized (resample to a common rate after proper anti-alias filtering, restrict to the common hardware passband, common reference, common channel subset) and state what cannot (noise floors, cap geometry). - Model setup explicitly (as a…
Age, sex or medication state differ between groups
- Symptom
- Patients older and sex-skewed relative to controls; patients recorded on medication.
- Cause
Clinical groups are rarely sampled the way controls are. Patients are older on average, drawn from a clinic rather than a community, more often of one sex for many conditions, and recorded while treated. Each of those variables has its own, well-documented effect on the EEG, and when it is unbalanced between groups the group comparison measures it too. The effect is compounded when the variable a…
- Detect
- Tabulate age, sex, medication state, handedness and any other covariate per group, with means, spreads and counts; report the imbalance rather than assuming matching. - Test whether the candidate biomarker correlates with the covariates within the control group; a measure that tracks age in controls will track an age imbalance between groups. - Compare artifact rates and data quality between gr…
- Fix
- Design: age- and sex-matched recruitment; record patients off medication where clinically possible, or record both states; document the covariates in participants.tsv. - Analysis: model the covariates (age, sex, medication) rather than only matching on them, with enough data to estimate their effects (L6.2); report the group effect with and without adjustment. - Interpretation: state the confou…
In other tools
In other toolsEEGLAB · FieldTrip — names only
The equivalents of what this lesson does, for a reader who works in another toolbox. Function names only: their own documentation is the place to learn how to call them.
EEGLAB
pop_spectopoEEGLAB
FieldTrip
ft_freqanalysis(mtmfft)FieldTripft_globalmeanfieldFieldTrip
Names checked 2026-09-18 against EEGLAB 2026.0.0 (plugins at the versions in EEGLAB’s own plugin list) and FieldTrip 20251218.
Reading
- Donoghue et al. (2020). Parameterizing neural power spectra. unverified
- Niso et al. (2022). Open and reproducible neuroimaging. unverified