Final capstone — Full reproduction
3 claims on this page are unverified. TODO(confirm) marks a specific statement the author has not yet checked against a
primary source. Everything else on this page has been reviewed. Treat a marked claim as
provisional and go to the cited source rather than quoting the sentence.
Brief
Choose a published EEG paper whose data are open. Reproduce its primary result, end to end, from the raw recordings: a preregistered analysis plan committed before you analyse anything, a C2-style pipeline, QC reports, the statistics, the figures, and a written report with a reproducibility statement. Publish the whole thing as a repository that someone else can run.
This is the last thing the ladder asks for, and it is the only capstone whose deliverable is a repository rather than a document. Everything below it was practice for one sentence: here is the analysis, here is the environment it ran in, here is what came out, and here is what you would have to do to disagree with me.
Four things make it harder than it sounds.
“From raw” means from raw. Not from the authors’ cleaned derivatives, not from a preprocessed release. The pipeline is the deliverable, and the point of starting from the recordings is that every choice between the file and the number becomes visible — which is exactly what C2 and L2.8 built you for. If the dataset ships a preprocessed version, you may use it as a comparison, never as your input.
Be fair to the paper. You can attempt this only because its authors published their data, and that is a selection effect running in their disfavour: the papers you can check are the ones whose authors made checking possible. A report that reads as an indictment has failed, whatever it found. Three habits keep it honest — distinguish “the paper does not state this” from “the authors did not do this”; read the supplement, the shared code and the data descriptor before recording an omission; and, where you had to guess a parameter, say that you guessed rather than implying the paper was silent by negligence. Most of what you will find is ordinary.
A divergence is a result, not a catch. If your numbers match, that is a finding about the analysis. If they do not, that is also a finding, and the work is to rank the candidate explanations — subset, exclusions, preprocessing, measurement definition, statistical model, software version, and ordinary numerical difference — by how much of the gap each could account for. What you may not do is tune your analysis until the two agree and then present the tuned version as the pre-registered one. Report the pre-registered result, then report the exploration separately and label it exploration.
“Runs from a clean environment” is a claim that has to be tested. Not asserted, not believed. Clone your own repository into an empty directory on a machine that holds none of your caches, build the environment from the file in the repository, run the entry point, and write down every single thing you had to supply that the repository did not state (L7.8). Whatever is on that list is the work that remains.
Reproduction is not replication, and the difference bounds every claim you may make
Say this in the report, because the two words are used interchangeably and they license different conclusions.
- Reproduction — the same data, analysed again. It tests the analysis: whether the reported numbers follow from the reported method applied to the reported data. A successful reproduction says the arithmetic and the description are sound. It says nothing new about whether the effect is real, because no new observations entered.
- Replication — new data, same question. It tests the finding.
You are doing the first. So a reproduction that agrees supports the paper’s reporting; a reproduction that disagrees points at the reporting, the pipeline or your own implementation, in that order of likelihood, and not at the phenomenon. “The effect is not real” is a claim your design cannot reach, and writing it would be the rubric’s fourth line failing in the most direct way possible.
How this differs from C6
C6 audits a paper against a reporting checklist and preregisters one analysis from it. C7 reproduces the primary result, whole, and publishes the machinery. If you have done C6 on the paper you choose here, reuse it: its audit becomes this report’s section on what the paper states, and its preregistration becomes the plan for the analysis you are reproducing, extended to cover the whole pipeline rather than one test. If you have not, you do not need to do C6 first — but read L6.6, because you will need its checklist to find out what the paper actually specified.
Choosing the paper
The starting list is the class-A directory entries — open and permissively licensed — that have a published primary result rather than a data descriptor.
The worked choice is ds-iowapd: Anjum et al. (2024), npj Parkinson’s Disease 10, 6 (doi 10.1038/s41531-023-00602-0), data at OpenNeuro ds004584 under CC0, openly downloadable, BIDS. It has the three things this capstone needs — a substantive claim, a real cohort, and data you can actually re-analyse — and it is the same dataset C6 uses, so the two capstones compose.
ds-dortmund, ds-srm, ds-pearl-neuro and ds-aszed are data descriptors. They document a resource rather than test a hypothesis, so there is no primary analysis to reproduce; choosing one turns this capstone into a resource description, which is a different and easier exercise. You may substitute any other paper whose data are genuinely open and whose primary analysis you can re-run — the requirements are a claim, a method section, and access.
Check the licence before you commit, not after
This is a capstone about publishing, so the licence of your input is a design constraint rather than a formality. What you may compute and what you may publish are different questions, and the directory records the answer for every entry:
- CC0 or CC BY (
ds-iowapd,ds-eegbci,ds-srm,ds-dortmund,ds-aszed,ds-brain-invaders,ds-areeg): derivatives may be published, with attribution where the licence asks for it. This is the class to pick from. - Share-alike (
ds-erpcore,ds-hbn, both CC BY-SA 4.0): derivatives you publish carry the same licence. That is a decision about your own outputs, and it has to be taken before you build a pipeline whose products you intend to release. - No-derivatives (
ds-bci-iv-2a, CC BY-ND 4.0): you may download it and compute on it, and you may not publish anything derived from it. A capstone repository whose deliverable is derived data cannot use it. - Non-commercial (
ds-mpeng,ds-chbmp): restricted to non-commercial use, and inds-mpeng’s case downstream of what participants consented to. Not a licence term anyone downstream can trade away. - DUA or registration (
ds-tdbrain,ds-brainlat): the terms are agreed by you and cannot be passed on, so a reader cannot re-run your work without agreeing them too. If you use one of these, the reproducibility statement has to say so plainly, and that is a real limitation on the deliverable rather than a footnote.
Record the licence in the repository, in the provenance, and in the report. “I checked the licence of every input before the first job ran” is a sentence very few published pipelines can honestly write.
What this site deliberately does not give you
This site does not reproduce the paper’s methods, its numbers or its figures anywhere. TODO(confirm): read the paper and the data descriptor yourself — everything in your report must be quoted or cited from them, and the dataset entry on this site is a catalogue record of the resource, not a summary of the analysis. If a number you need is not in the paper, that absence is itself a finding for the report, and the way to record it is “the paper does not state X, so I assumed Y, and here is what the result does under the alternatives”.
Inputs
- The paper. Read it before you read your own notes on it. You need the methods section, the supplement, the data-availability statement, and any code the authors shared. Read the shared code even if you do not use it: it is usually the only place the parameters that are missing from the methods actually appear.
- The data. For the worked choice:
ds-iowapd, “Rest eyes open” — OpenNeurods004584v1.0.0, dataset DOI10.18112/openneuro.ds004584.v1.0.0, CC0, access open, BIDS. 100 people with Parkinson’s disease and 49 controls; one eyes-open resting recording each, roughly three minutes; a 64-channel active cap on a DC amplifier at 500 Hz; Pz as the online reference and flat as a channel, with Iz, I1 and I2 excluded, so 60 channels are analysable; a 0.1 Hz online high-pass; 60 Hz mains; EEGLAB.set. The cohort’s documented confounds are part of the material and belong in the report: patients were recorded on dopaminergic medication, which attenuates beta signatures; the sex split is 68 male / 32 female among patients against 26 / 23 among controls; the groups differ in mean age; and the group sizes are unequal. See/datasets/iowapdand L7.6. - Your pipeline. The
pipelines/configuration from C2. Every parameter you change from that baseline is named in the report, with its reason and with the paper’s stated value beside it. - The checklist. L6.6. You are using it twice: to find out what the paper specified, and to make sure your own report specifies it.
- The preregistration structure. L6.4.
- The machinery. L7.8 for BIDS, containers, batch execution, the seven provenance items and the reproducibility statement.
Subset rule and runtime
The capstone notebook runs on a documented subset inside the ten-minute limit of §11 on Colab’s free tier, and scales to the full cohort locally:
- The documented subset is 10–20 participants per group, listed by ID in the notebook’s first cells and chosen by a rule that cannot depend on the result — the first n IDs in each group, or every participant passing the pre-specified exclusion rule in ID order. State the rule before you state the IDs, in that order, in the report.
- On the number itself. §11 sets the documented subset at 10–20 subjects, and it is written for a single-cohort capstone. This one contrasts two groups, where that figure has to be read per group or the design collapses. The binding constraint is the one §11 actually enforces — ten minutes on a free Colab tier — and the subject count is a proxy for it.
TODO(confirm): whether §11’s figure is meant per group for a two-group capstone is the author’s to settle; C6 and C7 read it the same way in the meantime. - A
FULL_COHORTswitch (defaultFalse) extends the run to the whole cohort locally. The per-participant function is identical; only the number of rows in the group table changes. Every number you quote says which state produced it. - Download only the participants you need, cache locally, delete raw downloads once the per-participant derivatives exist, and check free disk before you start.
- Per-participant derivatives that are expensive to recompute may be precomputed and shipped, so the group figures redraw without the full download.
- No absolute paths; every analysis parameter in one place with the seed; the final cell prints every number the report quotes.
A subset is not a limitation to apologise for — it is the thing that makes the repository runnable by a stranger, which is the rubric’s first line. What you must not do is quote a subset number and a full-cohort number in the same sentence without labelling both.
Deliverables
- The repository. A README that states in its first paragraph what this reproduces and what the result was; a licence for your code; a licence and attribution file for any derived data you publish; the directory layout; and one documented entry point that runs the analysis.
- The preregistration, reproduced verbatim in the report, with its commit hash and timestamp, committed before any group-level number exists. It covers the whole reproduction, not one test: the target result quoted from the paper, the subset and selection rule, exclusions with numeric criteria, the pipeline with every parameter, the measurement with its provenance, the model with unit of analysis, correction, α and tails, the covariates and why, what would count as a successful reproduction before you see your answer, and what you will do if a step cannot be run as written.
- The environment, as a container digest or a lock file, plus the record of the clean-environment test: the machine, the date, the commands, and everything you had to supply that the repository did not state.
- The data record. Dataset DOI and version, the exact subset with IDs, the selection rule, the licence, the download size, what you deleted afterwards, and the checksum of anything you ship.
- The pipeline, as configuration rather than prose, with every parameter that differs from the paper’s stated method listed in a table: your value, the paper’s value, and why. Where the paper does not state a value, say so and list the alternatives you tried.
- QC, reported by group. Channels flagged and interpolated, components removed with reasons, the fraction of data rejected, filter settings, trial or segment counts retained — per participant, aggregated per group. A cleaning step that behaves differently in patients than in controls is a group difference you manufactured, so this table is evidence and not housekeeping.
- The reproduction of the primary result, run as pre-registered: the effect in the units of the measurement, with a confidence interval, the standardised effect size with its variant named, the test in full, and the non-significant parts reported too.
- The figures, regenerated by your code from your derivatives — never traced from the paper — with units on every axis, uncertainty drawn, per-participant data visible where the sample allows, and the number of participants in every caption.
- The comparison table. The published quantity beside yours, each with the number of participants behind it and the measurement definition that produced it, so a reader can see whether the two numbers are the same quantity. Where the definitions differ, say so instead of putting the numbers side by side.
- The confound section this cohort demands. Age, sex, medication state and unequal group sizes (L7.6). Show the result with and without covariates, state plainly that adjustment is a modelling assumption rather than a repair, and say what design would settle it. If any acquisition variable differs between groups, name it as a site/device confound.
- A robustness report. A small multiverse over the two or three decisions most likely to matter — reference, filter cutoff, epoch length, artifact criterion — as a distribution with your pre-registered path marked (L6.4). Note that the reference is already constrained here: these data were recorded against Pz, so “no re-referencing” is itself a choice with consequences.
- The divergence paragraph. If your result differs from the paper’s, the candidate explanations ranked by how much of the gap each could account for, and what would distinguish between them. If it does not differ, say what your reproduction did and did not establish — the answer is not “the finding is confirmed”.
- The reproducibility statement. Inputs with DOIs, versions and licences; code repository and commit; environment digest; configuration hash; seeds; runtime and compute; units that failed and why; and — the part that makes it honest — what a reader cannot reproduce, and why.
- Your own limitations, in one paragraph: what you could not check, which of your judgements are contestable, and which parameters you had to guess.
Rubric
- It runs from a clean environment — the repository was cloned into an empty directory on a machine holding none of your caches, the environment was built from the file in the repository, the entry point ran, and the record of that test (machine, date, commands, everything you had to supply) is in the report.
- Every reporting-checklist item is present — the L6.6 checklist is worked against your own report, not only against the paper’s, and each item is stated in a form a stranger could act on.
- Divergences from the paper are explained — every parameter that differs from the paper’s stated method appears in a table with your value, theirs and the reason; every departure from your own pre-registered plan is listed as a departure, with the pre-registered version shown alongside.
- No claim exceeds the method — the report distinguishes reproduction from replication, does not conclude anything about the reality of the effect, and does not present post-hoc exploration as pre-registered analysis.
- The plan was committed first — the preregistration is in version control with a commit hash preceding every group-level result in the report, and the hash is quoted.
- The analysis starts from raw — no authors’ derivative is an input; any preprocessed release is used as a comparison and labelled as one.
- The subset rule is stated before the IDs, and every quoted number says whether it came from the subset or the full cohort.
- QC is reported by group, and any group difference in data quality or rejection rate is named rather than noticed later.
- The effect is reported in its own units with an interval, alongside a named standardised effect size; a p-value alone is not a result.
- The comparison table states both measurement definitions and both participant counts.
- The cohort’s confounds are addressed explicitly, and the limits of covariate adjustment are stated rather than assumed away.
- The robustness distribution shows the pre-registered path inside it, and is described as robustness rather than as a test.
- The treatment of the original paper is fair — “not stated” is distinguished from “not done”, the supplement and shared code were checked before any omission was recorded, guessed parameters are declared as guesses, and nothing is phrased as an accusation about the authors.
- Provenance is complete — input DOI and version, subset, code commit, environment digest, configuration hash, seeds, output checksums, and the list of units that failed with reasons.
- Every input’s licence is recorded, and it permits what you published.
- The reproducibility statement says what a reader cannot reproduce, and why.
- A reader who has the dataset could reproduce every number from the report alone.
Example report structure
- The paper and the claim. What it claims, in one paragraph, with the citation and the data DOI; which result is the primary one and how you decided; why you chose it.
- What the paper specifies, and what it does not. The methods as stated, item by item against the L6.6 checklist, with the parameters you had to supply marked as such.
- The preregistration, verbatim, with its commit hash and timestamp.
- Data. Dataset, version, licence, group sizes, subset rule, IDs,
FULL_COHORTstate, download size, and what was deleted. - Environment and the clean-environment test. Container digest or lock file, the test record, and the list of things you had to supply.
- Pipeline. Configuration, and the parameter table against the paper’s stated method.
- QC, by group, with the counts that survive each stage.
- The pre-registered result, in full: raw effect with an interval, standardised effect with its variant named, the test, and the non-significant parts.
- Confounds, with and without covariates, and what that comparison does and does not settle.
- Robustness, as a distribution with the pre-registered path marked.
- Comparison with the paper, with both measurement definitions and both Ns.
- Divergence, ranked — or, if there is none, what the agreement does and does not establish.
- Exploration, if any, clearly separated from everything above and labelled as exploration.
- Limitations, yours and the reproduction’s.
- Reproducibility statement, including what a reader cannot reproduce.
- One paragraph on what you would do differently if you were designing the original study — the last thing the ladder asks, and the only place in the report where you are allowed to speculate.
Estimated time: TODO(confirm) — the minutes value in the frontmatter is a site-design estimate for the whole capstone including the write-up, not a figure from the specification.
Rubric — self-assessment
Check each item you can honestly demonstrate in your report. This is the "submit" step: it is stored in your browser only.
0 / 4 rubric items checked. Self-assessment only; stored in this browser.