Neuroscience and Computational Methods
A route into neuroscience that runs alongside USMLE and clinical work. It starts at zero programming, uses EEG as a first practical laboratory rather than a permanent identity, gives you real exposure beyond EEG, and leads to one rigorous research capstone. A first-author manuscript is a serious conditional target, not a promised outcome.
eeg-lab.md (the EEG module) and templates.md (every form you copy from). Cross-references name the file and section, so each file works on its own.
Date-sensitive facts were checked on 31 July 2026. Programmes, meetings, dataset releases, and fees must be rechecked when you actually use them.
How to use this
Read Parts 0 to 4 once. Read Part 8 before your first session, because it governs what happens when a month goes wrong. Then work Part 5 in order. Parts 6, 7, 9 and the companion files are reference you consult, not material you complete.
Every block answers four questions: what is this, why now, what will I actually do, and what proves I finished.
Required
A test you attempt before studying. Pass it and you skip the instruction
Optional depth that never delays core work
The minimum that keeps the thread alive in a bad month
Cannot be tested out of, however fast you learn
A decision with a permitted stop, not a deadline
00 Start here: what neuroscience is, what you will actually do Part 0
0.1 What neuroscience actually is
Neuroscience is the scientific study of nervous systems. It asks how cells generate electrical signals, how cells connect into circuits, how circuits sense and control the body, how populations of neurons support perception and action, and how all of that relates to behaviour and disease.
It is not one subject and it is not another name for neurology. It connects several levels of explanation:
| Level | What is studied | Example question | Typical evidence | What that evidence cannot establish alone |
|---|---|---|---|---|
| Molecular | Channels, receptors, genes, signalling | How does a channel change excitability? | Molecular assays, genetics, pharmacology | Whole-circuit or behavioural consequences |
| Cellular | Neurons and glia | What makes a neuron spike? | Patch clamp, microscopy, cell manipulation | What a distributed network is doing |
| Circuit | Connected cell populations | How does excitation-inhibition balance shape a rhythm? | Multi-unit recording, calcium imaging, perturbation | A complete account of cognition or disease |
| Systems | Interacting regions and pathways | How is a movement planned and corrected? | Spikes, LFP, EEG/MEG, imaging, lesions, stimulation | Meaning or causality unless the design supports it |
| Cognitive/behavioural | Perception, learning, memory, decisions | How does evidence become a choice? | Psychophysics, computational behaviour models | The neural mechanism without neural evidence |
Clinical, computational, and engineering work cut across those levels rather than sitting above or below them:
| Lens | What it contributes | Example question | Main caution |
|---|---|---|---|
| Clinical | Frames dysfunction, diagnosis, prognosis, treatment | Which measurement changes a clinical decision? | Prediction does not establish mechanism or clinical utility |
| Computational | Equations, simulation, statistics, machine learning at any level | Which mechanism or representation reproduces the observation? | A model fit is not biological truth |
| Engineering | Instruments, interfaces, acquisition, stimulation | Can a device measure or alter the target reliably? | Technical performance is not scientific meaning |
Four questions to ask of everything you learn:
- What mechanism is proposed?
- What experiment or observation supports it?
- What exactly was measured?
- What conclusion remains unsupported?
That habit, not memorising subfields, is what starts making you a neuroscientist.
0.2 Neighbouring fields, without assigning you an identity
| Field | Central job | Relationship to this roadmap |
|---|---|---|
| Neurology | Diagnose and treat nervous-system disease | Your clinical background. Not the same as neuroscience |
| Clinical neurophysiology | Measure and interpret EEG, EMG, and related signals | A supervised clinical specialty. This roadmap gives research literacy, not certification |
| Cognitive neuroscience | Connect brain measurements to perception, memory, language, decisions | Sampled through systems material and non-EEG work |
| Systems neuroscience | Explain how circuits and populations generate function | Sampled through population data and mechanism questions |
| Computational neuroscience | Use formal models to understand nervous systems | A genuine strand here, both data-driven and theory-informed |
| Neuroinformatics | Organise, standardise, analyse, and share neural data | A strong methodological fit for the capstone |
| Neuroengineering | Build interfaces, devices, acquisition, stimulation | Adjacent. Not trained here |
| Clinical AI | Build or evaluate models meant to affect clinical decisions | One possible project family, not the definition of neuroscience |
You do not need a lifelong label in Block 1. You need a first laboratory, enough breadth to compare it against alternatives, and a defensible next decision.
0.3 What you will physically be doing
The work is less mysterious than the vocabulary. A session looks like one of these:
| Session type | What happens on screen | What remains afterwards |
|---|---|---|
| Mechanism | Read a section, redraw the causal chain, predict a perturbation | A mechanism map and a prediction |
| Coding | Open a notebook, change code, run it, read the error, fix it | A working notebook and an explained result |
| Data | Load a file, inspect metadata, plot raw values, check units and gaps | A data card and quality notes |
| Paper | Reconstruct question, design, measurement, figure, claim | A one-page figure audit |
| Question | Turn an observation into a card and try to kill it | A triaged, killed, parked, or promoted question |
| Mentor | Send one bounded artifact, ask one answerable question | Written feedback and a decision |
| Research | Run the locked analysis, log deviations, test robustness | Results and a decision log |
| Writing | Explain what you did, why, what happened, what stays uncertain | A section of the report |
Early on, most time is learning and producing small artifacts. After Gate A, the project becomes the vehicle through which you learn the next layer.
0.4 The terms you need before starting
Programming and tooling terms are explained in Part 5.2, immediately before you install anything.
0.5 What success means, and does not
A successful route ends with:
- a working map of neuroscience across levels and methods;
- the ability to connect mechanism, experiment, measurement, analysis, and inference;
- functional Python, statistics in code, validation, and neural-data skills;
- direct work with EEG, one population-spike example, and one decision-behaviour example, plus an explicit comparison against imaging;
- experience generating, killing, and refining research questions;
- one reproducible capstone, a full scientific report, and an honest decision about whether dissemination is justified;
- a lawful reproducibility archive and a next-question dossier;
- a real record of outreach, and ideally one sustained research relationship.
If the research track, Publication Release, and the authorship agreement all pass, the conditional outcome is a first-author manuscript, preprinted where appropriate and submitted to one journal. Acceptance is a later external event and is not required for this roadmap to have succeeded.
This roadmap does not promise a supervisor, programme admission, publishable novelty, journal acceptance, multiple papers, independent EEG interpretation, clinical certification, or a permanent specialty.
The paper is evidence that the apprenticeship worked. The apprenticeship is not reverse-engineered to manufacture a paper.
01 Direction, exploration, and the field map Part 1
1.1 The method spine and the first laboratory
Two things are fixed in advance. The scientific question is not.
Method spine: Python, quantitative reasoning, experimental design and inference, data handling and validation, and reproducible practice. These transfer to any neuroscience subfield.
First laboratory: EEG. Chosen because it is accessible without institutional gatekeeping, clinically legible to you, computationally tractable on a laptop, and makes acquisition, noise, labelling, and evaluation concrete.
The wording that keeps this honest:
EEG is my first working laboratory because it is accessible, clinically legible, and computationally tractable. It is not my professional identity. After structured exposure to other levels and modalities, I will explicitly choose whether to deepen, broaden, or switch.
If a strong mentor-led non-EEG opportunity appears, it may replace the EEG capstone through a written substitution memo (templates.md §11.10). It never runs in parallel as a second toolchain.
1.2 The field map you build
FIELD-MAP.md is a living table you add to from Block 1 onward. One row per topic you actually study:
| Column | What goes in it |
|---|---|
| Phenomenon | What is being explained |
| Level | Molecular, cellular, circuit, systems, cognitive |
| Lens | Clinical, computational, engineering, or none |
| Measurement | What was physically recorded |
| Design | Observational or experimental |
| Supported inference | What the evidence does establish |
| Unsupported claim | What it does not, however tempting |
| Your interest | High, medium, low, and why |
By the Interest Gate you should have 12 or more substantive rows spanning at least three levels and three measurement types. Twelve rows that span categories beat twenty that repeat one. That table, not a feeling, is what makes the Block 4 decision defensible.
1.3 The breadth floor
Before you narrow, you must have touched more than one kind of neural data. The floor is three thin slices, each using the same seven questions:
- What was measured?
- At what spatial and temporal scale?
- What preprocessing changed the observation?
- What is the unit of independence?
- What question can this answer?
- What causal claim can it not support?
- Did the scientific question interest you, or only the tool?
| Slice | Modality | Where |
|---|---|---|
| 1 | EEG time series | Block 2 |
| 2 | Population spikes | Block 3 |
| 3 | Decision behaviour plus a prepared fMRI-BOLD derivative | Block 4 |
These use small curated datasets and prepared notebooks. They are comparisons of measurement and inference, not invitations to install three preprocessing ecosystems.
1.4 The Interest Gate, end of Block 4
The one gate that is about you rather than the science.
You need: 12+ category-spanning field-map rows, three completed thin slices, and an honest note on which sessions energised you and which drained you.
You decide: deepen EEG, broaden before choosing, or switch home base.
If you choose to broaden: complete one additional predefined thin slice using at most one floating slot, then make the home-base decision. Exploration is bounded; it does not continue indefinitely.
Passes when the decision is written in one paragraph, references specific rows, and names what you are giving up.
Switching is not failure. Discovering in month 4 that population coding interests you more than clinical signals is the gate doing its job. What fails is drifting without deciding.
02 How research questions emerge and survive Part 2
No study is assigned in this document. Finding and killing questions is the transferable skill; a specific question is perishable.
2.1 Progressive question cards
Curiosities stay cheap. Only finalists earn expensive audit work. Full template in templates.md §11.3.
Observation, where it came from, level and domain, why it might matter, what you already know, what it would teach you, status.
Likely population and data, who benefits, obvious access or ethics or compute barrier, cost of the missing learning, the smallest test that could kill it, the fatal risk.
Precise estimand and comparator, five closest works and your differentiator, full feasibility check, intended-use statement, learning contract.
Statuses: curiosity, triaged, audit candidate, killed, parked, selected. Killing questions is progress. A log with nothing killed means nothing was examined.
2.2 Where questions come from
| Source | How to work it | Caution |
|---|---|---|
| Limitations and future-work sections | Read the last two paragraphs of twenty papers, log every stated gap | Authors' suggestions may be unimportant, infeasible, or already done. Treat as prompts, not endorsements |
| Failed reproduction | Try to reproduce a published result | Becomes research only when the failure is attributable, consequential, and documented against a precise claim |
| Benchmark and evaluation gaps | Where a community standard exists, non-standard evaluations can be redone standardly | Check the benchmark's own leaderboard first |
| Transportability | A model developed on one population, tested on another | Needs two genuinely compatible datasets |
| Label and reference-standard quality | What the ground truth is, how reliably it was made, what its noise does to measured accuracy | Needs annotation documentation to exist |
| Practice-literature mismatch | Something your service does that the literature does not describe | Uniquely yours, but needs access and usually ethics |
| Error analysis | The cases your own model gets wrong | Only available after you have built something |
| Mentor conversation | People with more questions than hands | Highest yield of all, once the relationship exists |
2.3 The dated overlap audit
Run before committing to a question. Produces a dated memo (templates.md §11.4).
Search all of these:
| Source | Why |
|---|---|
| PubMed | Biomedical baseline. Never sufficient alone |
| IEEE Xplore | Much neural-signal ML is engineering literature PubMed does not index |
| Scopus or Web of Science | Cross-disciplinary, via EKB |
| medRxiv, bioRxiv, arXiv | Current work appears here first, often a year before journals |
| Google Scholar forward citations | Take the closest paper, read everything citing it |
| PROSPERO, OSF | Registered but unpublished work |
| Conference proceedings | Frequently ahead of journals |
| Field benchmark pages | If a benchmark exists, the obvious comparisons are done |
The differentiator test. If you cannot state in one sentence what your study does that the closest existing work does not, you have an interest, not a project.
Audits go stale. Refresh briefly before protocol lock and again before any release decision. A timestamp records what you planned and when. It does not freeze the literature.
2.4 Feasibility, intended use, one estimand
Feasibility screen. Any failing row kills or reshapes the question.
| Dimension | Fails if |
|---|---|
| Data access | Needs a dataset whose approval you have not started |
| Independent units | Too few genuinely independent units to support the claim |
| Precision | No archetype-appropriate precision or feasibility assessment at the independent-unit level: expected confidence-interval width, minimum detectable effect, or a simulation-based justification. If precision is inadequate, narrow the estimand or claim, change the design, or stop the research route |
| Compute | Needs more than a laptop plus free tiers |
| Time | Cannot reach a frozen analysis inside the remaining blocks |
| Skills | Needs a technique costing more than the learning contract allows |
| Review | Needs domain adjudication nobody has agreed to give |
| Audience | You cannot name who would care |
Intended-use statement. One page before any modelling: population, setting, who would use the output, what decision it informs, the reference standard, what error is acceptable, the comparator, the primary endpoint. Without it, "good performance" has no meaning.
One estimand. The commonest way a first project dies is three primary questions. Pick one. Demote the rest to secondary or future work.
2.5 The project-learning contract
The mechanism that lets the project teach you the next layer without becoming an unfunded detour. Written at Gate A, naming exactly four things:
- One scientific concept to deepen.
- One method to learn.
- One implementation skill to acquire.
- One thing explicitly out of scope.
Then the loop, for each gap: classify it (concept, measurement, statistics, implementation, or interpretation), pick one primary resource and one exercise, predict what it should change before applying it, demonstrate it on toy or development data, integrate it, and write a short note on what it taught you.
If the missing layer cannot be learned and demonstrated inside Block 6, you narrow the question, choose a better-supported project, or park it. "Learning it is the point" is valid only when the time is actually there.
2.6 Gate A, end of Block 5
The most consequential decision in the roadmap.
Requires: 8 or more logged questions, all triaged; one full overlap memo (a second only for a genuine tie); a feasibility screen with no failing row; an intended-use statement; one named estimand; a learning contract; and one external reader who has commented on scope.
Three permitted outcomes:
| Outcome | When | What follows |
|---|---|---|
| Research project | A material, feasible question with a review route | Blocks 6 to 13 as the research track |
| Learning capstone | No question survives, but a rigorous bounded analysis exists | Blocks 6 to 13 as the structured learning track (Part 6.4). No novelty claim |
| Delay and redesign | Neither is ready | Use a floating slot, then reassess. Do not force a redundant study |
A learning capstone is a legitimate success. Manufacturing a novelty claim to protect a schedule is not.
03 Project directions, not assigned studies Part 3
Directions with worked shapes, not assignments. Your Gate A choice determines which template Blocks 6 to 13 follow.
| Archetype | Question shape | What it demands | Typical trap |
|---|---|---|---|
| Benchmark / re-evaluation | Do reported performances survive standardised re-evaluation? | Careful reimplementation, protocol discipline | Underestimating how long faithful reimplementation takes |
| Transportability | Does a model frozen on A hold on B? | Two compatible datasets, harmonisation reasoning | Datasets that look compatible and are not |
| Reproducibility audit | Can published pipelines be rerun at all? | Software forensics, patience, tact | A failure that is your environment, not their science |
| Label / reference-standard quality | How reliable is the ground truth, and what does its noise do? | Annotation documentation, measurement theory | Needing agreement data that was never published |
| Acquisition constraint | What survives cheaper, shorter, or lower-density acquisition? | Signal processing | A well-studied area where the answer is already known |
| Interpretability / human factors | Can a user read and act on the output? | Study design, ethics, participants | Needing human participants and approval you do not have |
| Mentor-supplied | Whatever the collaboration needs | A real role and a written substitution memo | Becoming unpaid labour with no authorship clarity |
Two worked examples of published work in the constraint family, useful as models of scope and method rather than as questions to take: Lin YC, Lin HA, Chang ML, Lin SF. Neurophysiol Clin. 2025;55(2):103044 (DOI); and Kojima J, Shi H, Ojemann WKS, et al. medRxiv preprint, 2026 (DOI). Read them to see what a finished question in this space looks like, and note that both occupy the montage-reduction question densely.
Human-factors, local-data, and clinically adjudicated projects are eligible only if ethics, participant or adjudicator access, and review support are secured by Gate A. Otherwise the first capstone stays inside a public-data template.
3.1 Evidence synthesis
A systematic review is not a lighter option. It is a team project that typically takes far longer than people expect.
In this roadmap: Block 13 permits only a scope, protocol, search, and throughput pilot for a possible later review. A full review can replace the empirical capstone only through a Gate A decision with a review-specific replacement map for Blocks 6 to 13.
Publication-grade synthesis requires independent duplicate screening and eligibility assessment, independent risk-of-bias assessment, checked extraction, and a design-appropriate conduct standard: Cochrane DTA and PRISMA-DTA for accuracy questions, QUADAS-3 for appraisal. PRISMA is a reporting guideline, not a conduct manual.
04 Data, measurement, and validity Part 4
4.1 A dataset is a scientific object
Access is part of the science, not an administrative preliminary. Every dataset you use gets a card (templates.md §11.5) recording version, provenance, licence, annotation origin, and known limitations.
chb24 was added later, bringing the collection to 198 seizures. chb21 is the same person as chb01, recorded 1.5 years later, so split at the person level, not the case level. chb24 is absent from the published subject-information file. If its identity cannot be resolved, exclude it from patient-independent confirmatory evaluation, or restrict it to exploratory and descriptive analysis, and say which you did. Channels are already bipolar derivations, with duplicated and reversed entries in some recordsrsync. Local governance checked separately. Request early4.2 Non-negotiable validity rules
These hold whatever your Gate A choice turns out to be.
- Name the exact corpus and version in every claim.
- Distinguish electrode potential, channel or derivation, acquisition reference, and display montage. Also distinguish membrane voltage from scalp voltage. Conflating these produces experiments you did not intend.
- Split on the unit of independence, which is usually the patient or session, never the window.
- Fit every learned transform inside training folds. Scaling, imputation, and feature selection included.
- Do not treat outer-fold scores as independent replicates.
- Final evaluation uses an untouched group-aware holdout, an external set, or prespecified nested resampling. A single fixed test set is one option, not a requirement.
- Count independent units before claiming precision. Nested observations inflate apparent sample size.
- Report metrics that carry the clinical or scientific meaning. For continuous detection that means event-level sensitivity and false detections per unit time, with any window-level summary clearly secondary.
- Record every deviation from protocol with its reason.
4.3 The leakage demonstration
Block 4's core exercise, and the single most valuable notebook in the roadmap. Train the same model three ways on the same data: split by row, split by group, and nested group-aware. Report all three numbers and explain the gap in writing. Once you have seen the inflation with your own data, you will never trust an unqualified accuracy figure again.
4.4 Reporting and appraisal chooser
Pick by design, do not stack all of them.
| Your design | Use |
|---|---|
| Prediction model, development or validation | TRIPOD+AI. Collins GS, Moons KGM, Dhiman P, et al. BMJ. 2024;385:e078378 |
| Diagnostic accuracy against a reference standard | STARD, and STARD-AI where AI is central |
| Appraising other people's prediction models | PROBAST+AI |
| Observational clinical study | STROBE |
| Systematic review | PRISMA and the design-specific extension |
| Anything else | Search EQUATOR before writing, not after |
4.5 The minimum reproducibility package
Every capstone ships with:
- pinned environment specification;
- dataset and version manifest with identifiers and, where licences permit, hashes;
- seeds and deterministic settings where feasible, with declared tolerances where exact equality is not achievable;
DECISIONS.mdlogging every analysis choice and its reason;- a README with a one-command reproduction path;
- a lawful archive, typically a versioned Zenodo deposit with a DOI;
- a small distributable or synthetic fixture so anyone can smoke-test the code path without holding restricted data.
For gated data such as TUSZ, reproduction means: another eligible and authorised user can reproduce the result from the documented lawful access and data-placement starting point, without undocumented assistance from you. You cannot ship the data; you can ship everything else.
The archive is a reproducibility artifact, not a publication. It does not count as an output and no roadmap in this file should imply otherwise.
5.1Calendar logic
Thirteen ordered blocks. Fifteen calendar slots. Two slots hold no mandatory work.
There is no hour ledger and no reserve bank. Time cannot be saved in March and spent in September, so the plan does not pretend otherwise. What protects you is spare calendar, not accumulated credit.
| Rule | Detail |
|---|---|
| One at a time | Only one core block is open. Finish it or formally pause it |
| Blocks are ordered | Each depends on the last. Reordering breaks the dependencies |
| Load labels, not hours | Light, Standard, or Full. Sized to fit roughly one to two hours a day, clustered however suits you |
| Slots absorb disruption | A block spilling into the next slot is the system working |
| Rebaseline, never catch up | See Part 8 |
The completion rule, canonical.
All Core items are required unless explicitly marked remediable, movable, optional, or conditional. The done-when test adds the required integration or transfer demonstration; it does not replace the Core list.
A block ends when its Core items are done and its done-when test passes, not when its content has been read. Block N+1 does not start until Block N passes. That is what stops overfill turning into invisible debt.
When the target date moves: only when remaining blocks exceed remaining slots. Not automatically after a bad month. Check at the end of any compressed or paused slot.
5.2Tools, explained before you install them
Every tool below exists to solve a specific problem. Read this before Block 1.
Python is a programming language. You will use it to load data, compute things, and make figures. You do not need computer-science depth; you need enough to express an analysis and read an error message.
A package is code someone else wrote that you import. numpy for arrays and maths, pandas for tables, matplotlib for plots, scipy for statistics, scikit-learn for models, mne for EEG. You will almost never write these things yourself.
An environment is an isolated set of packages with pinned versions. It exists because Project A needing version 1.2 and Project B needing version 2.0 will otherwise break each other, and because "it worked on my machine" is not reproducible. Conda creates and manages environments; Miniforge is a lightweight installer for it. When you export an environment.yml, you are recording exactly what your results depended on.
A notebook is a document mixing code, output, figures, and prose, run in chunks. Good for exploring and explaining. Its weakness is that cells can be run out of order, so a notebook can appear to work while being irreproducible. Hence the habit, from Block 1 onward, of restart the kernel and run from the top before you believe anything.
Git tracks changes to files over time. A commit is a saved snapshot with a message saying what changed and why. It exists so you can see what you did three weeks ago, undo a change that broke something, and prove when work was done. GitHub hosts Git repositories online, giving you backup, a link you can send someone, and a public record. A repository or repo is one project's tracked folder.
Zotero stores papers and generates citations. ORCID is a free permanent researcher identifier that journals and data providers ask for.
The exact sequence. Run these in order. If a step fails, fix it before the next one.
# 1. after installing Miniforge, create the environment
conda create -n neuro python=3.11 -y
conda activate neuro
# 2. install everything this roadmap uses
conda install -c conda-forge numpy pandas matplotlib scipy scikit-learn \
jupyterlab mne -y
# 3. record exactly what you have, and commit this file
conda env export --from-history > environment.yml
# 4. set up version control in your project folder
git init
git add .
git commit -m "Initial environment and smoke test"
--from-history records only what you asked for rather than every transitive dependency, which makes the file portable across operating systems.
What you do in Block 1:
Do not reinstall a working setup to match this list. Never put non-public papers or any patient material in a public repository.
5.3The roadmap at a glance
| # | Block | Load | Gate | Ends with |
|---|---|---|---|---|
| 1 | Orientation, baseline, working system | Standard | Environment runs; baseline recorded; people mapped | |
| 2 | Python foundations, mechanisms, first EEG contact | Full | Real EEG file loaded and described correctly | |
| 3 | Statistics in code, circuits, spike data | Full | Estimate with uncertainty; population thin slice | |
| 4 | Validation, systems, cross-modal, Interest Gate | Full | Interest | Leakage demonstrated; home base chosen |
| 5 | Question triage and Gate A | Standard | A | One question selected and defended, or learning track |
| 6 | Project-specific learning and development pilot | Full | The next layer learned and demonstrated | |
| 7 | Development workflow and draft protocol | Standard | Pipeline runs end to end on development data | |
| 8 | External review, separation, Gate B | Standard | B | Protocol locked and registered before results |
| 9 | Primary analysis | Full | The prespecified result exists | |
| 10 | Robustness, error analysis, claim control | Full | Result survives or fails prespecified stress | |
| 11 | Freeze, clean rerun, Gate C | Standard | C | Everything reproduces from a clean environment |
| 12 | Full scientific report and Release decision | Full | Release | Report complete; dissemination decided honestly |
| 13 | Execute the decision, close out, then stretch | Standard / Full | Archived always; submitted only if Release passed; next-question dossier | |
| Two floating slots | Absorb disruption. No mandatory content |
5.4Your first thirty days
The only week-by-week section, so that "start" has a visible meaning. It covers Block 1.
If setup eats the month, use a floating slot. Do not delete the orientation or pretend the environment works.
01
Orientation, baseline, and a working system
LoadStandard
FocusSetup, scope, baseline, people
GateNone
Compressed
Paused
0%0/0 done
You are building the small operating system through which everything else becomes visible and recoverable. The aim is not to learn Python or survey neuroscience in a month. It is to understand the kind of science you are entering, find out what you can already do, and make the next action runnable. Setup has lead time you cannot compress later, which is why access requests and people-mapping start now rather than when you need them.
Before touching the setup instruction, try to: create and activate an isolated Python environment; load a small CSV, inspect types and missingness, compute a grouped summary, make a labelled plot; initialise version control, commit, restart, rerun; and explain in plain language what the rows, code, and plot represent. If that works unaided on unfamiliar data, skip the installation tutorials and the Kaggle modules. The smoke test, the written explanation, and the reproducibility record still stand: they are evidence, not instruction.
The smoke test runs from a restarted kernel, your four tracking files exist and have real content, the Block 1 transfer test passes (templates.md §11.11), and you can explain what an environment and a commit are without looking them up.
Add two more field-map rows in an area you know nothing about.
Finish the smoke test and one outreach row. Nothing else. Everything here is a dependency for Block 2.
02
Python foundations, mechanisms, and first EEG contact
LoadFull
FocusData handling, neuron mechanism, real EEG
GateNone
Compressed
Paused
0%0/0 done
Real EEG appears now rather than in month six. Early contact surfaces file-format, naming, memory, and metadata problems while they are still cheap. It also makes the mechanism reading concrete: you will read about membrane voltage and then look at scalp voltage, and the difference between those two is the point.
On an unfamiliar tabular dataset, load it, inspect types and missingness, compute a grouped summary, produce a labelled figure, and explain what one row represents. If clean, compress the pandas instruction but not the EEG work.
You can pass the Block 2 table and montage cases in templates.md §11.11 without the key, and explain out loud how electrode potential, channel or derivation, display montage, membrane voltage, and scalp voltage differ. No waveform diagnosis required or permitted.
Open CHB-MIT file chb01_01.edf, recognise that names like FP1-F7 are already bipolar derivations, reconstruct two of the equations, and explain why ordinary re-referencing is not recoverable from them. Do not apply an average reference to derived channels.
Read one mechanism subsection, keep five retrieval prompts alive, do one outreach action. Resume the EEG work before Block 3.
03
Statistics in code, circuits, and spike data
LoadFull
FocusEstimation with uncertainty, population data
GateNone
Compressed
Paused
0%0/0 done
Three things usually taught separately come together: circuit and population reasoning, a quantitative estimate with honest uncertainty, and a real neural observation. You learn one analysis family properly rather than surveying a catalogue of tests. This is also the second thin slice, so you now have two modalities to compare.
On unfamiliar paired data, fit a regression and generate a bootstrap interval. Explain the sampling unit, what the interval represents, why a narrower interval is not necessarily less biased, and how grouped observations would change the resampling. If sound, skip the exposition but not the coded transfer.
You can run and interpret the same estimate-with-uncertainty family on the fixed transfer case, reuse your unchanged raster function on the withheld seeded unit, and state the spike subset's preparation, selections, session nesting, and generalisation limit. Before excluding the top 1% of units by spike count, predict whether mean or median moves more, then check. Your three normal EEG descriptions are complete and use the form.
Run only the first demonstration in Neural Rate Models and predict one parameter effect. Do not start dimensionality reduction or deep learning.
Finish or repair the figure audit and record the exact next notebook cell. Do not open a second analysis family.
04
Validation, systems, cross-modal inference, and the Interest Gate
LoadFull
FocusLeakage, modality comparison, home-base decision
GateInterest Gate
Compressed
Paused
0%0/0 done
Two things at once. You learn the most common computational validity failure, which is letting information from the same person or session sit on both sides of an evaluation. And you complete the breadth floor, so the Interest Gate is a decision made on evidence rather than on mood. Everything after this block is shaped by what you decide here.
Given a grouped dataset, identify the correct unit of splitting, build a leak-free pipeline with all transforms inside the folds, and explain why an apparently excellent score can be an artefact. If clean, compress the tutorial and go straight to the demonstration.
The boundary. This is exposure to an analysis-ready fMRI derivative, not training in raw MRI preprocessing. Do not add FSL, fMRIPrep, registration, segmentation, another course, or a predetermined imaging project at this stage.
The visual-circuit reading has moved to Block 8’s reviewer-waiting period so this block does not grow.
The leakage demonstration is complete and explained, all three thin slices are done, your six artefact cards each name a prevention, confirmation, and handling strategy, and the home-base decision is written and defensible.
Repeat the leakage demonstration on a second dataset with a different grouping structure.
Complete the leakage demonstration only. It is a dependency for everything after Gate A. The gate paragraph can slip into a floating slot.
05
Question triage and Gate A
LoadStandard
FocusOverlap audit, feasibility, selection
GateGate A
Compressed
Paused
0%0/0 done
This block decides whether the next eight are worth doing. Deliberately lighter on new content, because auditing is slow, and because a question that survives real scrutiny is worth more than one that arrives fast. Most of the work is reading and killing your own ideas.
None. Question selection is a research act and cannot be tested out of.
Gate A has one of its three permitted outcomes recorded in writing (Part 2.6), with the memo, screen, intended use, estimand, and contract attached. The EEG transfer test has been passed on the first attempt or on the single retest. If the retest also failed, the interpretation dependency is recorded in writing and expert adjudication is required for all interpretation-sensitive claims from here on.
Draft the first version of the protocol.
Keep auditing the leading candidate. Do not select a question you have not audited because a slot is ending.
06
Project-specific learning and a development-only pilot
LoadFull
FocusThe next layer, on your own project
GateNone
Compressed
Paused
0%0/0 done
This is where the project starts teaching you. The learning contract from Gate A names exactly what you need, and this block is when you get it: not a survey, but the specific concept, method, and implementation skill your question requires. The pilot exists to find out what breaks before anything expensive is built on top of it.
Attempt the contract's method on toy data before studying it. Anything you can already demonstrate compresses to a check rather than a course.
The contract skills are demonstrated on data you did not practise on, the pilot runs end to end, and every result you have already seen is written down so it can be disclosed at Gate B.
Start the protocol draft.
Finish the contract demonstration. The pilot can move; unlearned methods cannot.
07
Development workflow and draft protocol
LoadStandard
FocusReproducible pipeline, protocol draft
GateNone
Compressed
Paused
0%0/0 done
You turn a working notebook into a workflow someone else could run, and you write down what you intend to do before you find out what happens. Both are what separate a result from an anecdote.
If you already write config-driven, restartable pipelines with logging, skip the workflow instruction and go to the protocol.
The pipeline runs from configuration alone on a development subject you did not build it on, the protocol draft names every analysis you intend to run, and a named reviewer has agreed to read it.
Identify and approach your external reviewer for Block 8.
Keep DECISIONS.md current. It is the artifact that is hardest to reconstruct later.
08
External review, separation, and Gate B
LoadStandard
FocusReview, registration, lock
GateGate B
Compressed
Paused
0%0/0 done
Locking a protocol before you see the primary result is what makes the result mean something. It is also the point where an outside reader can still change the design cheaply. After this block, changes are deviations to be declared rather than choices to be made.
None. External review and registration are research acts.
The protocol is registered, the pilot disclosure is written, and you have not looked at a primary result. A registration proves what you planned and when. It does not prove novelty.
A dynamical-systems or domain-theory tutorial.
Wait for review rather than locking early. Gate B is worth a slot.
09
Primary analysis
LoadFull
FocusExecute exactly what was locked
GateNone
Compressed
Paused
0%0/0 done
Run what you prespecified, in the order you prespecified it. The discipline here is the entire value of the previous block. Every deviation goes in the log with its reason, and deviations are not failures as long as they are declared.
None. This is execution.
You can state the primary result in one sentence including its uncertainty, and every departure from the protocol is written down.
Begin the robustness analyses that were prespecified in Block 7.
Run the primary analysis only. Robustness moves to Block 10 where it belongs anyway.
10
Robustness, error analysis, and claim control
LoadFull
FocusStress the result, bound the claim
GateNone
Compressed
Paused
0%0/0 done
A single number is not a finding. What makes it credible is showing how it moves under the perturbations you named in advance, and looking honestly at the cases it gets wrong. This is also where the claim gets sized to the evidence rather than to your hopes.
The result has survived or failed its prespecified stress tests, the error sample has been examined, and your claim is sized to your independent units rather than your window count.
Begin the report skeleton.
Run the prespecified sensitivities. Skip exploratory work entirely.
11
Freeze, clean rerun, and Gate C
LoadStandard
FocusReproducibility
GateGate C
Compressed
Paused
0%0/0 done
Work that does not rerun is the commonest reason competent research is rejected. If it does not reproduce for you in a clean environment, it will not reproduce for anyone else.
None. Reproduction is a research act.
Someone else could clone the repository and regenerate your figures without asking you a question.
Start the report.
The rerun is the block. Everything else waits.
12
Full scientific report and the Release decision
LoadFull
FocusWrite it up, then decide honestly
GatePublication Release
Compressed
Paused
0%0/0 done
The report is the deliverable, and it exists whether or not anything is ever submitted. Writing it is also how you find out what you actually understand. The release decision comes after the report, not before, so that the writing is not bent toward a predetermined conclusion.
None. Writing is a research act.
The report is complete and the release decision is written down with reasons.
Begin manuscript conversion.
Finish the report. The release decision can move to Block 13.
13
Execute the decision, close out, then consider a stretch
LoadStandard
FocusShip, archive, and plan forward
GateNone
Compressed
Paused
0%0/0 done
Two jobs. Convert the decision into an action, and leave the work in a state your future self can pick up. Stretch work is considered only after both are done and only if the calendar genuinely allows.
The release decision has been acted on, the archive exists, and the next-question dossier is written.
Archive and dossier only. The stretch is the first thing to go and should go without regret.
6.1 Two tracks, both successful
| Track | When | Endpoint |
|---|---|---|
| Research capstone | Gate A found a material, feasible question with a review route | Full report, plus manuscript if Release passes |
| Structured learning capstone | No question survived, but a rigorous bounded analysis exists | Full report, explicitly making no novelty claim |
Both produce a complete scientific report, a reproducibility archive, and a next-question dossier. The learning capstone omits only the novelty claim and the manuscript. It is a real outcome and should be described accurately rather than apologetically.
6.2 Capstone Completion is not Publication Release
Two separate decisions. Conflating them is how educational analyses end up impersonating publishable research.
Capstone Completion asks: is the work rigorous, reproducible, honestly reported, and complete? If yes, the apprenticeship succeeded.
Publication Release asks all of the following, and every one must hold:
| Criterion | Fails if |
|---|---|
| Material contribution | A refreshed audit shows the question is answered |
| Sufficient evidence | Independent units cannot support the claim |
| External scientific review | Nobody qualified has examined the design and interpretation |
| Reproducibility | Gate C did not pass cleanly |
| Lawful release | Data agreements or ethics do not permit it |
| Authorship agreement | Contributors have not agreed roles and order |
If Completion passes and Release fails, the roadmap succeeded. Say so plainly.
6.3 Accurate states
Four different things, never collapsed:
| State | Means |
|---|---|
| In preparation | Being written |
| Submitted | Sent to a venue, decision pending |
| Accepted | Editorial acceptance received |
| Published | Publicly available in final form |
A preprint has a persistent identifier and is not peer reviewed. A Zenodo archive is a reproducibility artifact, not a paper. Journal acceptance and indexing are external events on someone else's timetable.
On software: the Journal of Open Source Software requires software developed privately to have at least six months of public development history before submission, plus releases, public issues, tests, documentation, and feature-complete scope; single-purpose analysis pipelines and thin utilities are explicitly out of scope. A capstone pipeline built inside this roadmap will not qualify. Archive it on Zenodo, cite it, and stop there.
6.4 The structured learning track
If Gate A selects this route, Blocks 6 to 13 keep the same shape with three substitutions: the question becomes a well-posed analytic exercise on public data; the overlap audit becomes a positioning statement acknowledging existing work; and the report states explicitly that no novelty claim is made. Everything else, including protocol lock, prespecified analysis, robustness, and reproduction, stays identical. The rigour is the point.
6.5 Report rubric
| Section | Must contain |
|---|---|
| Question | The estimand in one sentence, and why it matters |
| Background | What is known, what is not, where this sits |
| Data | Corpus, version, provenance, units, exclusions, limitations |
| Methods | Everything needed to rerun, plus the protocol reference and every deviation |
| Results | The primary estimate with clustering-aware uncertainty, then secondary and clearly labelled exploratory analyses |
| Robustness | Prespecified sensitivities and the error analysis |
| Limitations | Reference-standard quality, independence, generalisation, what would change the conclusion |
| Conclusion | Sized to the evidence, naming the nearest thing the data do not support |
| Reproducibility | Environment, manifest, seeds, tolerances, archive location |
6.6 Authorship
Agree roles and order in writing before analysis, against the ICMJE criteria. Authorship disputes on student-led work are common and almost entirely preventable. If someone reviews your protocol, be explicit about whether that is a courtesy, a supervisory role, or authorship.
7.1 Mentorship is a workstream, not a wish
Gate A requires an external reader. You cannot gate on a relationship you have not built, so outreach starts in Block 1 and continues throughout.
| Block | Outreach action |
|---|---|
| 1 | Map 8 to 12 people. Profiles for all; read properly for the best three |
| 2 | Send the first three messages: one bounded artifact, one answerable question |
| 3 | Share the figure audit with anyone who replied |
| 4 | Second wave, informed by your Interest Gate decision |
| 5 | Ask your best contact to comment on Gate A scope |
| 8 | Request protocol review |
| 12 | Request report or manuscript review |
| 13 | Update the map with every relationship that now exists |
What works: a specific named paper of theirs, a specific artifact of yours, one answerable question, and a bounded ask. What does not: generic mass mail, unbounded requests for supervision, or five attachments.
Silence is data, not a pass. Log non-responses. Silence documents that you tried; it does not satisfy a gate on the research track. If nobody qualified responds by Block 5, your permitted options are: approach alternates, spend a floating slot, narrow the question and its claims to what unreviewed work can support, or take the learning track. A knowledgeable colleague’s comments are peer feedback and must be labelled as such, not as expert review.
7.2 Access is scientific feasibility
Data access belongs in the feasibility screen, not in an administrative appendix. Requests with approval steps start in Block 1 if they might be needed. A dataset you cannot get is not a limitation to work around; it is a question you cannot ask yet.
7.3 Programmes substitute, they never stack
Structured programmes are accelerators, not dependencies. All Neuromatch materials are open, so the curriculum is yours regardless of admission. Arabs in Neuroscience runs an Arabic-language computational neuroscience course with open materials and an in-person school via the IBRO portal; its paid teaching-assistant role, which carries a certificate and a possible reference letter, is the most valuable thing either organisation offers you, and your existing biostatistics teaching makes you a credible candidate.
If admitted to a live intensive, that period becomes programme-only. The programme's curriculum replaces planned curriculum, its group project replaces planned project work, and the affected blocks move. It is never additional. Write the substitution memo (templates.md §11.10).
Dates, fees, formats, and application windows change. Check the official page when the action arrives, never this document.
7.4 Conferences
Build a venue table from official calls when results make dissemination plausible, which is Block 12 or later, not as a monthly ritual from Block 1. Never copy a historical deadline forward as fact. Confirm presenter requirements before submitting, since several major meetings require onsite presentation and an abstract you cannot present is a smaller asset than it appears.
Given that results arrive late in the roadmap, the realistic first presentation is a regional meeting after the plan ends. That is normal and worth saying rather than pretending otherwise. Only submit to meetings run by an actual professional society or university.
8.1 Four modes
Triggered by circumstances, not scheduled in advance. You do not need to know your exam dates for this to work.
| Mode | When | What you do |
|---|---|---|
| Normal | Ordinary weeks | The block as written |
| Compressed | Dedicated exam study, heavy rotation, or life | Reading, retrieval, and writing, plus finishing or repairing an already-defined technical workflow. Never a new toolchain, a new exploratory branch, or a high-stakes analysis. If conditions do not permit careful execution, defer the task rather than doing it badly. See the block's if compressed line |
| Paused | No capacity | Nothing. Log the date in PROGRESS.md |
| Re-entry | First one or two sessions back | Read PROGRESS.md, rerun the last command that worked, do the smallest next artifact. Do not start something new |
8.2 The two floating slots
Two of the fifteen calendar slots hold no mandatory content. They exist to be consumed by disruption. Using them is the system working as designed, not falling behind.
8.3 Rebaseline, never catch up
After any compressed or paused slot, do one thing: count remaining blocks against remaining slots.
- Blocks fit: continue. Change nothing.
- Blocks exceed slots: either move the target date or drop optional scope. State which, in writing.
Never stack extra hours to recover a slot. Never compress a research act to protect a date.
8.4 Fast learning: demonstrate, then compress
Speed is handled by evidence rather than assertion. Every block opens with a diagnostic. Pass it and the instruction compresses or disappears.
Time saved goes to depth first: a harder transfer case, an extra figure audit, more cross-modal exposure, or better project work. It does not automatically become another output.
8.5 What cannot be tested out of
These are not instruction, so no diagnostic can skip them:
- inspecting the real data yourself;
- locking a protocol before results;
- validation and leakage control;
- debugging;
- reproduction from a clean environment;
- obtaining external review;
- writing the report.
Learning speed shortens the time to understand a method. It does not shorten data access, reviewer response, collaborator decisions, or journal processes.
8.6 Substitutions and maintenance
Links rot, datasets change, programmes move. When a resource has changed at the moment you need it: record what changed, find a lawful equivalent serving the same learning job, and note the substitution. Never silently swap a dataset, and always record the version you actually used.
8.7 Project stop rules
| Trigger | Action |
|---|---|
| No question survives Gate A | Take the learning capstone track. Do not force a novelty claim |
| Independent units cannot support the estimand | Narrow the claim or stop at Gate A or B |
| The pilot shows the analysis will not fit | Reshape at Gate A rather than discovering it at Block 10 |
| Access fails after Gate A | Substitute a lawful dataset, or switch archetypes with a written memo |
| The result is null | Report it. A well-conducted null with a locked protocol is a real finding. Nonsignificance does not establish absence: claiming a negligible or absent effect requires a prespecified equivalence-type claim or an interval that excludes effects you would consider meaningfully large |
| Six months pass with no shipped artifact | Stop learning and force a small complete piece of work |
| Risk | What absorbs it |
|---|---|
| No reviewer responds | Log the silence, use peer feedback with the limitation named, or change route. Silence never passes a gate by default |
| Access route changes or closes | Recheck at use, record the version, substitute lawfully. Never silently swap |
| No material question survives Gate A | Named rigorous learning capstone without a novelty claim |
| Sample cannot support the estimand | Narrow the claim, or stop at Gate A or B |
| Result is null or non-novel | Preserve protocol and report. Release may fail while the apprenticeship succeeds |
| Coauthor or journal timing exceeds the calendar | Record the accurate state and extend. Never call preparation "submitted" |
| USMLE or clinical work consumes capacity | Switch mode, consume a floating slot, or move the date. No catch-up debt |
| Learning runs faster than planned | Diagnostics compress instruction; saved time goes to depth, not to extra outputs |
| Tutorial hell | Every block ends in an artifact plus a completion test appropriate to its stage: a transfer demonstration early, a locked decision or reproduction later. Six months without a shipped artifact triggers the stop rule |
| Scope creep after Gate B | Only prespecified sensitivities are confirmatory. Everything else is labelled exploratory |
| One project consumes the whole calendar | Stop rules and the Gate A feasibility screen exist precisely to prevent this |
Neuroscience. Neuroscience Online, the open-access electronic textbook from the Department of Neurobiology and Anatomy, McGovern Medical School at UTHealth Houston. Neuromatch Computational Neuroscience, open curriculum.
EEG. See eeg-lab.md for the full source list, built on the American Epilepsy Society atlas on NCBI Bookshelf and the ILAE 2025 seizure classification.
Standards. EQUATOR Network hosts all reporting guidelines. TRIPOD+AI: Collins GS, Moons KGM, Dhiman P, et al. BMJ. 2024;385:e078378. STARD-AI. PROBAST+AI. QUADAS-3.
Data. PhysioNet · CHB-MIT · TUH EEG including TUSZ · OpenNeuro · IBL
Tools. MNE-Python · scikit-learn · Zenodo · OSF
Signal processing. Mike X Cohen's free lectures. One conditional paid course, decided at Gate A: see eeg-lab.md section 3.
Worked examples of finished questions in the EEG constraint family. Lin YC, Lin HA, Chang ML, Lin SF. Neurophysiol Clin. 2025;55(2):103044. DOI. Kojima J, Shi H, Ojemann WKS, et al. medRxiv preprint, 2026. DOI.
All facts checked 31 July 2026. Recheck anything date-sensitive when you use it.