Why reproducibility matters in evidence synthesis
Systematic reviews and meta-analyses are used to make decisions about care, policy and funding, so errors in them travel far. Studies that re-extracted data from published meta-analyses have found discrepancies, and some have found that a corrected analysis changed the conclusion. A review that cannot be checked is one whose errors will not be found, at least not before they matter.
Reproducibility also pays for itself. A review with scripted analysis and organized data can be updated by rerunning the code on new studies, which turns a months-long project into a matter of weeks. It supports living systematic reviews, secondary analyses and teaching. And it protects the authors: if a reader questions a figure, the answer is a file, not a recollection.
Two terms are worth distinguishing. A result is reproducible if the same data and code give the same numbers. It is replicable if an independent team, following the same protocol with new searching and extraction, reaches a similar conclusion. This guide is mainly about the first, which is within the control of the review team, and which makes the second easier to test.
Four layers of a reproducible review
It helps to think of the review in four layers, each of which should be recoverable.
- The question and plan. The protocol, with its registration, and a dated record of any changes.
- The search and selection. The strategies as run, the dates, the exported records, the deduplication log, the screening decisions and the reasons for exclusion.
- The data. The extracted dataset, the source of each value, the risk-of-bias judgments and any derived values with the calculation.
- The analysis. Code that reads the dataset and produces every table and figure, with the software environment recorded.
A failure at any layer breaks the chain. A pooled estimate can only be reproduced if the data are available, and the data only mean something if they can be traced to studies that were found by a documented search.
Organizing and sharing the data
The most valuable file is the extracted dataset. Keep it in a plain, open format such as CSV, with one row per study-outcome-comparison and clearly named columns. A data dictionary should define every column, its units and its coding, and should give the allowed values. Keep a column for the source of each key number (report, table, page) and a flag for values that were calculated, estimated from figures or supplied by authors. Store the original extraction forms from each extractor, and the consensus version, if double extraction was used.
Sharing the dataset is now expected by many journals and by open-science policies. Deposit it in a repository that gives a persistent identifier, such as the Open Science Framework, Zenodo or an institutional repository. The data in a review come from published sources, so copyright and privacy issues are usually small, but consider any individual participant data separately, since those need data-sharing agreements and may not be shareable. If some data cannot be shared, say so and say why.
Scripted analysis
Analyses done by hand, by clicking through menus, cannot be repeated exactly. Scripted analysis in R, Stata, Python or a similar language can. R is common in meta-analysis, with the metafor, meta and netmeta packages, and Stata has built-in commands and user-written ones. A minimal analysis in R with metafor may look like this:
library(metafor)
dat <- read.csv("extracted_data.csv")
dat <- escalc(measure = "RR", ai = events_t, bi = nonevents_t,
ci = events_c, di = nonevents_c, data = dat)
res <- rma(yi, vi, data = dat, method = "REML", test = "knha")
summary(res)
predict(res) # includes the prediction interval
forest(res)
funnel(res)
The example reads the dataset, computes log risk ratios, fits a random-effects model with the restricted maximum likelihood estimator and the Knapp-Hartung adjustment, and produces the main outputs. Good scripts go further. They create every table and figure that appears in the manuscript, so that nothing is copied by hand. They set a random seed wherever randomness is involved, as in Bayesian sampling and bootstrapping. They contain comments that explain choices. They run from top to bottom without manual steps.
Recording the software environment
Results can change between software versions, for example when a default estimator or a rounding rule changes. Record the version of the language and every package used. In R, the function sessionInfo() prints it, and tools such as renv can freeze the package versions in a lock file. In Python, a requirements file or a conda environment does the same. For long-term preservation, containers such as Docker can package the whole environment. Stata users should record the version and any user-written commands. This is a small effort at the time of analysis and an enormous help to anyone, including you, who tries to rerun the work years later.
Version control and naming
Use version control, such as Git, for the code and, where practical, for the data. It records who changed what and when, makes it possible to return to an earlier state and allows collaboration without overwriting. Platforms such as GitHub, GitLab and OSF provide hosting. Tag the version that produced the published results, so that it remains identifiable when development continues. For files that are not suitable for version control, adopt a consistent naming convention with dates and never overwrite a file that has been used for a result.
Reproducing the search
Search results change over time, because databases are updated and indexing is revised, so a search cannot be repeated exactly later. What can be reproduced is the strategy and what it found at the time. Save the strategy for each database exactly as run, the date, the interface and the number of records, and export the full result set in a standard format such as RIS. PRISMA-S sets out the reporting items for searches. Keep the screening records, with the reviewer decisions and reasons, in a file that can be shared. Together these allow another team to see precisely which records were considered and why studies were excluded.
Protocols, registration and deviations
An open protocol is the first reproducibility document, because it fixes the plan before the results are known. Registration with PROSPERO, the Open Science Framework or another registry gives it a date and an identifier. Deviations from the protocol happen in nearly every review, and they are not a problem if they are recorded. Keep a dated log of changes with the reasons, and report them in the paper. When analyses were added after seeing the data, label them as post hoc. This protects the review from the suspicion that outcomes or analyses were selected to give a particular answer.
What a replication package contains
- The registered protocol and a log of deviations.
- The full search strategies, dates and the exported records or their identifiers.
- The screening decisions and the list of excluded studies with reasons.
- The extraction forms and data dictionary.
- The final analytic dataset, with sources for key values.
- The risk-of-bias assessments.
- The analysis code and a description of how to run it, usually a README file.
- The software environment: versions, lock files or a container.
- The outputs: tables and figures as produced by the code.
- A licence, so that others know how they may reuse the material.
The package should be deposited in a repository with a persistent identifier, and cited in the paper in the data and code availability statement. Test it by asking a colleague who did not work on the review to run it on a clean computer. The problems they find are the problems a future reader would meet.
Limits and cautions
Reproducibility has limits. Data that came from individual participants may be restricted. Some software is proprietary and not everyone has it. Judgments, such as risk-of-bias ratings and eligibility decisions, cannot be reproduced by code, only documented. And a reproducible review is not necessarily a correct one: it may reproduce the same mistake faithfully. Reproducibility makes errors discoverable, which is its value, but it does not replace careful design, double extraction and critical appraisal.
A workable routine for a review team
Reproducibility is easiest when it is built into the routine from the start, not retrofitted at submission. A practical routine has five habits. First, create a project folder with a fixed structure at the start: a place for the protocol, for search exports, for screening files, for extraction forms, for the analytic data, for code and for outputs. Second, agree on file names and on who owns each file. Third, write the data dictionary when the extraction form is designed, not afterwards. Fourth, write the analysis script as soon as the first studies are extracted, using the incomplete data, so that bugs are found early and the final run is a matter of pressing one button. Fifth, keep a short dated log of decisions, covering eligibility rulings, handling of ambiguous data and changes to the plan.
The log is the most underrated item. A line such as "2026-03-14: studies reporting only change scores are included and analysed separately, agreed by both reviewers" costs seconds to write and answers a question that otherwise takes hours to reconstruct, or cannot be answered at all.
When data cannot be shared
Some materials cannot be shared in full. Individual participant data usually come with agreements that forbid onward sharing. Commercial databases may restrict redistribution of exported records. In such cases, share what is permitted, such as the identifiers of records rather than full citations, aggregate data, or synthetic data with the same structure so that the code can be tested, and describe how a qualified researcher can request access. State the restriction in the data availability statement, with the reason. A partial package that is honest about its limits is far more useful than none, and it shows that sharing was considered and not merely avoided.
How we can help
We can deliver analyses as documented scripts with data dictionaries, record environments, organize data and search records for deposit, and assemble a replication package alongside the manuscript. [OWNER VERIFICATION REQUIRED] The relevant services are statistical analysis, data extraction and coding and protocol development and registration.
Frequently asked questions
What makes a meta-analysis reproducible?
Shared extracted data with a data dictionary, scripted analysis that produces every number, a recorded software environment and documented search and selection.
Do I have to share my data?
Many journals and funders expect it. The extracted data from published studies are usually shareable. Individual participant data may need agreements, and restrictions should be stated.
Which software is best for reproducibility?
Any scriptable software works. R with metafor, meta or netmeta is common, as is Stata. What matters is that the analysis can be rerun from a script.
Can the search be reproduced?
Not exactly, as databases change. Save the strategy, dates, interface and exported records so that what was found at the time can be seen.
Where should I deposit the materials?
In a repository that provides a persistent identifier, such as OSF, Zenodo or an institutional repository, and cite it in the data and code availability statement.
What is the difference between reproducible and replicable?
Reproducible means the same data and code give the same results. Replicable means an independent team reaches a similar conclusion with new searching and extraction.
References
- Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1-48.
- Balduzzi S, Rucker G, Schwarzer G. How to perform a meta-analysis with R: a practical tutorial. Evid Based Ment Health. 2019;22(4):153-160.
- Gotzsche PC, Hrobjartsson A, Maric K, Tendal B. Data extraction errors in meta-analyses that use standardized mean differences. JAMA. 2007;298(4):430-437.
- Lakens D, Hilgard J, Staaks J. On the reproducibility of meta-analyses: six practical recommendations. BMC Psychol. 2016;4:24.
- Wilkinson MD, Dumontier M, Aalbersberg IJJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018.
- Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews. Syst Rev. 2021;10:39.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71