Evidence synthesis in neuroscience
Neuroscience covers a wide range of methods and scales. Synthesis takes different forms accordingly. In human imaging, coordinate-based meta-analysis looks for locations where findings from many studies converge. In cognitive and behavioral work, ordinary meta-analysis of effect sizes combines performance differences between groups or conditions. In animal research, systematic reviews gather evidence on mechanisms and treatments before clinical trials. Clinical aspects are treated on the neurology page.
The field has faced a reproducibility discussion. Studies are often small, analysis pipelines are flexible, and results are published selectively. Synthesis can help, since pooling reduces some noise, but it cannot remove biases that affect every study. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for life sciences.
Coordinate-based neuroimaging meta-analysis
| Approach | Input | Notes |
|---|---|---|
| Activation likelihood estimation (ALE) | Peak coordinates from each study | Tests spatial convergence of reported peaks; does not use effect sizes; requires correction for multiple comparisons |
| Seed-based d mapping | Peak coordinates plus statistics, with optional unthresholded maps | Uses effect size information; can model signed differences |
| Multilevel kernel density analysis | Peak coordinates | Treats contrasts as the unit; different weighting |
| Image-based meta-analysis | Full statistical maps | Most accurate; needs data sharing |
| Region-of-interest meta-analysis | Effect sizes in prespecified regions | Standard meta-analysis; requires prespecified regions |
Most imaging papers report only the coordinates of statistically significant peaks, and not full maps or effect sizes, so many meta-analyses use methods that test whether peaks cluster in the same places more than chance. Results depend on the inclusion criteria, the threshold for significance used in the primary studies, the software and the correction for multiple comparisons. A review should state each, register the protocol, report the number of contrasts and subjects per analysis, and check that a few studies do not drive the result. Studies that used a whole-brain analysis should be separated from those restricted to regions of interest, since the latter bias the spatial distribution of peaks.
Statistical power and small samples
Analyses of neuroscience studies have found that median statistical power is low, so many true effects are missed and significant findings tend to overestimate effect sizes. In imaging, sample sizes of around 20 per group are common, though large datasets are now available. For meta-analysis this has two consequences. Pooled estimates from many underpowered studies can still be informative, but they are vulnerable to selective reporting, which is most severe when power is low. And effect sizes from small studies are less likely to hold in large preregistered samples. A review should compare effects in small and large studies and report the numbers of participants contributing to each analysis.
Reverse inference and interpretation
Finding that a brain region is active during a task does not show that the region is specific to a mental process, because many regions respond to many tasks. Inferring a mental state from activation in a region is called reverse inference, and it is only as strong as the specificity of the region for the process. Meta-analytic databases can estimate how selective a region's response is across tasks, which helps, but caution is still needed. A review should describe findings as patterns of convergence across studies and avoid language that equates a region with a function. Structural differences between groups, such as a smaller volume in a disorder, are associations that do not reveal their causes or show diagnostic use.
Animal studies and mechanisms
Preclinical neuroscience uses rodents and other species to study mechanisms, drugs and disease models. Systematic reviews and meta-analyses of such work have been used for stroke, pain, neurodegeneration and psychiatric models. They have shown that reporting of randomization, blinding and sample size calculation is often poor and that studies without these report larger effects. The usual effect size is the standardized mean difference, and heterogeneity is large because of differences in species, strain, sex, model and dosing. Female animals have been under-used, and a review should report the sex of animals and compare results where it can. Risk-of-bias tools for animal studies and the ARRIVE guideline apply here.
Cognitive and behavioral measures
Cognitive neuroscience meta-analyses combine performance on tasks across groups, for example differences between patients and controls or effects of brain stimulation. Tasks differ in their names and in what they measure, and effects often depend on task parameters. Reliability of cognitive tasks can be low, which limits correlations with brain measures and the stability of group differences. A review should record the task, the measure derived from it, its reliability if reported and the population. Brain stimulation studies also need attention to sham controls and blinding, since participants may perceive real stimulation.
Electrophysiology and event-related measures
Studies of EEG and MEG report amplitudes and latencies of evoked responses and oscillatory power. Meta-analyses of such measures need to group together comparable components, electrodes, reference choices and processing steps. Effect sizes are standardized differences in amplitude, and small samples make them imprecise. Methodological choices, such as filtering and artifact rejection, affect results and are rarely reported in full. We code what is reported and describe missing information as a limitation.
Publication bias and analytic flexibility
Imaging studies have many analytic choices, and exploration can produce significant findings by chance. In coordinate-based meta-analysis, tests such as the fail-safe N approach for ALE show how many null studies would be needed to remove an effect, though they have been criticized. Where possible we check for small-study effects, examine registered reports and preregistered studies separately, and encourage use of shared data. The field is moving toward data sharing, which allows better syntheses in the future.
Connectivity, structure and clinical comparisons
Many neuroimaging syntheses compare patient groups with controls on gray matter volume, white matter integrity or functional connectivity. These studies face confounds that a review should record: age, sex, medication, illness duration, head motion and scanner differences. Findings in disorders such as schizophrenia, depression or dementia often overlap across diagnoses, which has led to meta-analyses that test for shared and specific abnormalities. Effects are usually small to moderate and vary with sample characteristics. Studies of connectivity add methodological choices about preprocessing and network definitions that change results, so we report them and avoid combining analyses that used very different pipelines without a moderator.
A group difference in a brain measure is an association. It does not show that the difference caused the condition or that the measure could diagnose it in an individual. Diagnostic claims require studies of accuracy, which are assessed using the methods described on the diagnostic accuracy page, with attention to whether the accuracy was tested on people who were not part of the discovery sample. Reports of high accuracy from a single small sample, tested on the same people used to build the model, are a known source of overstated claims.
Common pitfalls we look for
- Mixing whole-brain and region-of-interest studies in coordinate-based analysis.
- Interpreting regional activation as proof of a mental process.
- Counting multiple contrasts from the same participants as independent.
- Ignoring low power and its effect on effect-size inflation.
- Omitting randomization and blinding details in animal studies.
- Using lenient thresholds in primary studies without noting them.
Planning a neuroscience synthesis
We help define the question and the modality (imaging, electrophysiology, behavior, animal studies), plan searches in MEDLINE, Embase, PsycINFO, Web of Science and bioRxiv, and set up coding of task, contrast, sample, threshold, software, correction method and animal characteristics. For coordinate-based analyses, we specify the software and settings in the protocol before extraction. See the meta-analysis service for scope and process.
An invented example of power and inflation
Suppose the true standardized difference between two groups is 0.3 and studies enroll 20 per group. The standard error of the difference is about the square root of 2 divided by 20, which is 0.32. Power to detect 0.3 at the 5 percent level is therefore low, roughly 16 percent. If only significant results are published, the studies that pass the filter must show an observed difference of at least about 0.62, more than double the true value. A meta-analysis of the published set would then report something near that exaggerated figure. This is why a review compares small with large studies and why, in a low-power field, pooled estimates from the published literature can mislead.
Coding and data management
Coding frames record modality, field strength or recording equipment, task and contrast, sample size, participant characteristics, medication, threshold and correction, software, coordinate space and animal features. Coordinates reported in different spaces are converted to a common space with documented methods. Two coders work independently on a sample, and the coded coordinates and code are shared with the final report so others can repeat the analysis.
Responsible interpretation
Findings about the brain attract media attention and are easily overinterpreted. A review explains what convergence of findings does and does not show, avoids claims about individuals and states limits plainly. The service does not provide clinical interpretation of scans or diagnoses.
How we support research projects in this area
From experiments to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A structured question, effect-size choice, plan for dependent effects and registration where a platform accepts it.
Searching and extraction
Searches across biological databases, extraction from text, tables and figures with checks between extractors.
Multilevel analysis
Multilevel, phylogenetic and meta-regression models, with publication-bias analysis adapted to dependent data.
Manuscript and submission
PRISMA-EcoEvo or PRISMA 2020 checklists, the manuscript, data and code for sharing.
Boundaries of this service
A neuroscience synthesis describes published findings on group differences, task effects and animal experiments. It does not interpret scans or recordings for individuals, provide diagnosis or recommend treatment. Small samples and flexible analysis make many findings uncertain, and associations between brain measures and behavior do not show cause.
Frequently asked questions
What is coordinate-based meta-analysis?
A method that tests whether reported activation peaks from many imaging studies converge on the same locations more than chance.
Why is statistical power a concern?
Low power means true effects are missed and significant results overestimate the effect, which affects pooled estimates from a selected literature.
Does activation in a region show a mental process?
Not by itself. This reverse inference depends on how selective the region is, so reviews describe convergence and avoid equating regions with functions.
How are animal studies assessed?
With risk-of-bias tools, attention to randomization, blinding, sex and model, and checks for publication bias.
Can imaging studies using different thresholds be combined?
With care. Thresholds and correction methods are recorded, and studies are checked for undue influence on results.
Do you interpret brain scans for patients?
No. The service provides research and evidence-synthesis support only.
References
- Button KS, Ioannidis JPA, Mokrysz C, et al. Power failure: why small sample size undermines the reliability of neuroscience. Nat Rev Neurosci. 2013;14(5):365-376.
- Eickhoff SB, Laird AR, Grefkes C, Wang LE, Zilles K, Fox PT. Coordinate-based activation likelihood estimation meta-analysis of neuroimaging data: a random-effects approach based on empirical estimates of spatial uncertainty. Hum Brain Mapp. 2009;30(9):2907-2926.
- Poldrack RA. Can cognitive processes be inferred from neuroimaging data? Trends Cogn Sci. 2006;10(2):59-63.
- Müller VI, Cieslik EC, Laird AR, et al. Ten simple rules for neuroimaging meta-analysis. Neurosci Biobehav Rev. 2018;84:151-161.
- Macleod MR, Michie S, Roberts I, et al. Biomedical research: increasing value, reducing waste. Lancet. 2014;383(9912):101-104.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.