Level 3 L3.3 judgment thread

Measuring ERPs

Peak versus mean versus area, latency measures, defining windows without circularity, and difference waves.

~60 min Widget: w-measurement-explorer Notebook: nb-3-3-measurement

Prerequisites: L3.2 · The ERP and its components

3 claims on this page are unverified. TODO(confirm) marks a specific statement the author has not yet checked against a primary source. Everything else on this page has been reviewed. Treat a marked claim as provisional and go to the cited source rather than quoting the sentence.

Objectives

  • Choose among peak amplitude, mean amplitude and area
  • Measure latency (peak, fractional area, jackknife)
  • Define measurement windows without circularity (a priori, collapsed localizer)
  • Compute difference waves

Why this matters

Everything after this lesson operates on numbers, not on waveforms: the statistics, the effect sizes, the comparison with the literature. How you turn a waveform into a number decides whether a difference between conditions is a difference in brain activity or a difference in how noisy the two averages were. Two of the most natural choices — pick the peak, and put the window where the effect looks biggest — are the two that break most reliably, and both break in the direction of finding an effect.

Concepts

Amplitude: peak, mean, area

Peak amplitude is the largest (or most negative) value inside a measurement window. It is intuitive, it matches a long literature, and it is biased. The bias has a simple source: taking a maximum is not a linear operation, so noise cannot cancel. Adding symmetric noise to a waveform can only ever push its maximum further out, never pull it back, and therefore the expected peak of a noisy average is larger than the peak of the underlying signal. Everything that changes the noise level of an average — trial count, rejection, filtering, electrode impedance, participant group — changes the measured peak amplitude in a direction. That is pf-peak-amplitude-noise-bias, and it is the mechanism by which condition-biased rejection (L2.5) manufactures effects.

Peak amplitude has a second, quieter problem: it depends on the width of the window. A wider window can only find a larger maximum, so two studies with different windows are not measuring the same quantity.

Mean amplitude is the average of the waveform over a fixed window. It is linear, and that single property is why it is the modern default. Linearity means the mean amplitude of an average equals the average of the mean amplitudes, so the measurement commutes with averaging across trials, with difference waves and with grand-averaging across subjects. It means noise makes the measurement imprecise but not biased: the expected value is right whatever the noise level, so two conditions with different trial counts can still be compared. And it means the value does not change when a wider window is used for a symmetric reason — it changes only because more of the waveform is included, which is a decision you can inspect.

Area comes in two forms and it is worth keeping them apart. Signed area (the integral over the window) is mean amplitude multiplied by the window length — the same statistic in different units. Rectified area, which integrates the absolute value or only the positive part, is not linear, and it inherits the bias problem of the peak: noise adds area that cannot cancel. Rectified area is useful when polarity varies across subjects and dangerous when noise varies across conditions.

The practical ranking for a condition comparison: mean amplitude over an a-priori window first; signed area when you want the same thing in area units; peak amplitude only when the literature you are speaking to demands it, and then with the trial counts and noise levels reported alongside.

Latency

Peak latency is the time of the extremum within the window. Everything said above about peak amplitude applies, and more: the latency of a noisy waveform jumps between local maxima rather than moving smoothly; it is undefined when the component appears only as a shoulder; and it is bounded by the window, so at low signal-to-noise it drifts toward the window centre.

Fractional area latency is the time at which a stated fraction of the area within the window has accumulated — conventionally 50%, giving a centre-of-mass measure. Every sample in the window contributes, so it moves smoothly as the data change and is far more stable at low signal-to-noise. It measures a different thing from the peak: a waveform that is skewed early has its 50% area latency before its peak, and that is not an error.

Fractional peak latency is the time at which the waveform first reaches a stated fraction of its own peak — commonly 50%, as an onset-like measure. It is more stable than the peak latency itself but still inherits the peak’s definition, so a noisy peak gives a noisy reference level.

Jackknife scoring attacks the problem from the other side. Instead of measuring each subject’s own noisy waveform, compute N grand averages, each leaving one subject out, and measure the latency on each of those — waveforms with roughly N times the trials, on which a peak or an onset is actually locatable. The leave-one-out measures vary far less than individual measures do, by a known amount, so the statistic computed from them must be corrected: the standard correction divides the resulting t statistic by (N − 1). What you give up is per-subject values — there is nothing to correlate with a behavioural measure, and nothing to put into a mixed model. TODO(confirm): the exact correction and the conditions under which it holds should be checked against the primary jackknife-scoring literature, which is not among the site’s reading-list keys.

Measurement windows, and circularity

A measurement window is a claim about where the effect is, made before the effect is measured. Three defensible ways to arrive at one:

  • A priori. Take the window from the published literature on the component, or from a previous experiment in your own lab, and fix it in the analysis plan. Simple, transparent and slightly lossy — your recording’s component may not sit exactly where the literature’s did.
  • Collapsed localizer. Average all conditions together, choose the window and the electrodes from that collapsed waveform, and then measure the condition difference in the window you chose. Under the null hypothesis the collapsed average is independent of the condition difference, so the selection does not bias the test. (Luck, 2017) is the standard treatment; it is the method to reach for when the component’s latency in your data is uncertain.
  • An independent localizer. Define the window from separate data — a different block, a different task, the other half of a split — which is the cleanest option and the most expensive.

What is not defensible is choosing the window where the difference looked biggest. The cost is not a matter of taste: you have implicitly searched over every window you would have accepted, so the effective number of tests is far greater than one, and the type I error rate rises accordingly. That is pf-post-hoc-windows, and L6.4 quantifies it. The same logic covers every other selection made with the effect in view: which electrodes, which measure, which subjects, which polarity.

Judgment call

Write down the measurement — component, window, electrodes, method, polarity — before you plot the condition difference, and keep the written version. The diagnostic question is not “was my window reasonable?” but “how many other windows would I have found reasonable?” If the answer is more than one, the analysis has degrees of freedom that the p-value does not know about, and the honest options are a collapsed localizer, a pre-registered window, or a method that tests the whole window at once (L3.7).

Difference waves

Subtract one condition’s average from the other, per subject, and measure on the difference. Three reasons this is usually the right object to measure:

  • Overlapping components cancel. Whatever both conditions share — the sensory response to the stimulus, the neighbouring components that distort a peak — subtracts out, leaving something closer to the latent component the design manipulated.
  • The reference largely cancels (L2.3). A difference wave travels between labs with different references far better than a raw amplitude.
  • The measured quantity becomes the tested quantity. You measure the effect, rather than measuring two things and subtracting afterwards, which matters for non-linear measures where the two are not the same.

The costs are real. Variances add, so the difference wave is noisier than either condition — with equal trial counts, by a factor of about √2. And a difference wave cannot tell you whether one condition rose or the other fell, which is often the interesting part. Standard practice is to plot both conditions and their difference, and to measure on the difference.

The lateralized components are difference waves by construction: the N2pc and the LRP are contralateral minus ipsilateral, computed within subject and within hemisphere. That construction is what makes them sensitive to eye movements, because a systematic saccade toward the target is itself a lateralized difference (pf-eye-movements-lateralized, L3.5).

The data behind this lesson

  • ds-erpcore P3, CC BY 4.0, open access, per-subject downloadable; 30 EEG + 3 EOG channels, 1024 Hz, CMS reference, 60 Hz mains, no software filters, 40 participants per paradigm. TODO(confirm): the author mirrors the ERP CORE entry into the catalogue registry and signs off the dataset page (§10.11 item 8); the shipped asset sidecars also record a licence conflict in the source — the OSF node record says CC BY 4.0, the per-paradigm component’s own LICENSE file says CC BY-SA 4.0 and its dataset_description.json says CC0 — which the author reconciles (§13 item 22).
  • The widget serves sub-001’s 200 single trials at Pz — every trial, including the ones a rejection criterion would drop, so the trial count can be thinned for real rather than simulated — together with per-condition average waveforms for 20 subjects on the same time grid, which is what lets the across-subject variance panel update as you change method and window. TODO(confirm): the per-trial rejection flags behind those subject averages are label_source: algorithmic until the author reviews them (§4.5).
  • The notebook measures the N170 and the P3 several ways and compares their reliability, so the amplitude argument above is checked on two components with very different shapes.

Explore

Measurement explorer — switch measurement method and window on the same waveform and watch the value and its across-subject variance move

mode: default Open lab page →
Loading Measurement explorer — switch measurement method and window on the same waveform and watch the value and its across-subject variance move…

What to look for

Data: ds-erpcore · license CC-BY-SA-4.0 · DOI TODO(confirm) · labels: algorithmic · a cropped, re-referenced or filtered derivative of the source recording. Share-alike. This asset is derived from a source whose licence requires that anything built from it carry the same licence. If you reuse it, distribute your version under CC-BY-SA-4.0 and keep the attribution below.
Kappenman, E., Farrens, J., Zhang, W., Stewart, A. X., & Luck, S. J. (2020). ERP CORE: An Open Resource for Human Event-related Potential Research. PsyArXiv. Paper DOI 10.31234/osf.io/4azqm. Dataset DOI 10.18112/openneuro.ds003069.v1.0.0, https://osf.io/thsqg/. Licensed CC-BY-SA-4.0; this is a derived asset and is distributed under the same licence (data/directory.yaml license_decision, section 13 item 15).

The sequence that makes the argument: with the noise slider at zero, switch between peak and mean amplitude and confirm that the two agree about which subject is large; raise the noise and watch every subject’s peak amplitude climb while the mean amplitudes stay put and merely scatter; open the across-subject panel and compare the spread of the two measures at the same noise level; then drag the window narrow and wide and watch the two measures respond differently. Finish with the latency measures: step the noise up and down while watching peak latency jump between local maxima and the 50% area latency move smoothly.

Practice

Measuring ERPs: peak, mean, area, peak latency and fractional-area latency compared for noise bias and across-subject reliability nb-3-3-measurement

Level 3 ~5 min
notebooks/L3/nb-3-3-measurement.ipynb

Downloads from ds-erpcore.

Open in Colab Download Read it here

The notebook measures the N170 and the P3 with peak amplitude, mean amplitude, signed area, peak latency and 50% fractional-area latency, on full and thinned trial sets, and reports the across-subject reliability of each. Its final cell prints the two numbers this lesson’s exercise asks for.

Exercises

Measure the P3 on the low-trial subject the notebook identifies, at Pz, in the a-priori window it states.

Exercise ex-3-3-peak-amplitude

Numeric

Peak amplitude of the P3 for the low-trial subject

µV

Accepted within ±0.3 µV.

Exercise ex-3-3-mean-amplitude

Numeric

Mean amplitude of the P3 for the same subject, same window

µV

Accepted within ±0.2 µV.

Exercise ex-3-3-explain-divergence

Free response

The two numbers differ by more for this subject than for the high-trial subjects. Explain why, and say which of the two you would carry into a group analysis.

Exercise ex-3-3-window-choice

Multiple choice

Your component's latency in this sample is uncertain, so an a-priori window from the literature may be misplaced. Which procedure for choosing the window keeps the condition test valid?

Options

Pitfalls

Pitfall

Peak amplitude inflated by noise

Symptom
Low-trial condition has larger peaks.
Cause

Peak amplitude is the maximum (or minimum) of the waveform inside a window, and taking a maximum is not a linear operation. Adding symmetric, zero-mean noise to a waveform can push its largest sample further out but can never pull it back, so the expected peak of a noisy average is larger than the peak of the underlying signal:

Detect
  • Report trial counts per subject and per condition, always, and look at them before looking at the effect. - Measure the same effect with mean amplitude over the same window. If the effect is present with the peak and absent with the mean, the peak’s bias is the leading explanation. - Plot the measurement against trial count across subjects. A downward-sloping relationship between trial count an…
Fix
  • Use mean amplitude over an a-priori window as the default measurement. It is linear, so its expected value does not depend on the noise level, and it commutes with averaging, difference waves and grand-averaging (L3.3). - Use signed area when area units are wanted; avoid rectified area when noise differs across the compared units. - For latency, use a fractional-area latency rather than a peak…

Full entry with example →

Pitfall

Windows chosen from the tested data

Symptom
Effect only in the window that looked biggest.
Cause

A measurement window, an electrode set, a measure and a polarity are all selections. Making a selection from the same data you then test means the test no longer has the error rate it claims, because the selection has already searched over alternatives the test knows nothing about.

Detect
  • Read the methods for where the window came from. “Determined from the grand average of the condition difference” is the failure; “determined from the grand average collapsed across conditions” is not. - Count the choices you would have accepted, not the ones you made. If more than one window, electrode set or measure would have been defensible, the analysis has degrees of freedom the p-value do…
Fix
  • Fix the window in advance, from published work or from your own previous data, and record it before unblinding. - Use a collapsed localizer when the component’s latency in your sample is genuinely uncertain: choose the window and electrodes from the average of all conditions together, which under the null carries no information about the condition difference, and then measure the difference the…

Full entry with example →

In other tools

In other toolsEEGLAB — names only

The equivalents of what this lesson does, for a reader who works in another toolbox. Function names only: their own documentation is the place to learn how to call them.

EEGLAB

  • pop_geterpvaluesERPLAB plugin (install separately)

Names checked 2026-09-18 against EEGLAB 2026.0.0 (plugins at the versions in EEGLAB’s own plugin list).

Reading

  1. Luck (2014). An Introduction to the Event-Related Potential Technique, 2nd ed.. unverified
  2. Luck & Gaspelin (2017). How to get significant effects in any ERP experiment. unverified