Guide

Data extraction in systematic reviews

Data extraction is where a review turns published papers into numbers, and it is where many errors are made. Extraction mistakes are common, they are rarely noticed, and they can change a pooled result. This guide covers form design, piloting, double extraction, the calculations that are often needed, and the record that should be kept.

Why extraction needs care

Extraction means reading each included report and recording study characteristics, participants, interventions, outcomes and results in a structured way. It looks clerical, but it involves many decisions: which of several reported outcomes to use, which time point, which analysis population, whether a number is a standard deviation or a standard error, and whether a value in a figure can be read. Empirical studies of systematic reviews have found error rates in extracted data that are not trivial, and that in some cases they changed the conclusion. Independent double extraction reduces these errors substantially compared with extraction by one person.

The extracted data are also the evidence base for everything that follows, including the forest plots, the heterogeneity statistics, the risk-of-bias judgments and the certainty ratings. Care here protects the whole review.

Designing the extraction form

A form is written before extraction begins and is organized by the questions the review must answer. It normally has four parts.

  • Study identification. Report identifiers, linked reports, country, funding source and conflicts of interest, and registration number.
  • Methods and risk of bias. Design, randomization and concealment, blinding, attrition, and the items required by the chosen risk-of-bias tool.
  • Participants and interventions. Inclusion criteria, setting, sample size, age, sex, condition severity, details of each arm including dose, duration and delivery, and the comparator.
  • Outcomes and results. For each outcome of interest: definition, scale and direction, time points, number analysed, and the effect data, either means and standard deviations or event counts, or an effect estimate with its interval, and the analysis model used.

Only extract what the review will use. A form with fifty fields that no analysis needs slows the work and increases the chance of mistakes. On the other hand, collect the characteristics that will be needed for subgroup analyses and meta-regression, because going back to papers later is costly. These should be planned in the protocol.

Use a structured tool. Spreadsheets are common, and dedicated systems such as Covidence, DistillerSR, EPPI-Reviewer and the Systematic Review Data Repository provide forms, dual extraction and consensus functions. Use data validation, drop-down lists and clear definitions of each field.

Piloting and double extraction

Test the form on a few studies of different types, preferably by every extractor, and compare. Fields that are interpreted differently need clearer instructions, and fields that nobody could fill should be dropped. Record the instructions in a data dictionary.

Then extract in duplicate. Two people extract independently, and a third resolves differences. A reduced approach, in which one person extracts and another checks, is faster and is accepted for some purposes, such as rapid reviews or for descriptive data, but studies of extraction suggest that independent double extraction of outcome data is the safer standard. Whatever approach is used, state it in the methods, and report the agreement or the number of discrepancies and how they were resolved.

Calculations that are often needed

Reports do not always give the data in the form the analysis needs. Several conversions are standard, and each should be recorded with the original numbers.

Standard deviation from a confidence interval. A report gives a mean with a 95 percent confidence interval of 18.0 to 22.0 in 50 participants. The standard error is the interval width divided by 3.92, and the standard deviation is the standard error multiplied by the square root of the sample size: sqrt(50) x (22.0 - 18.0) / 3.92 = 7.22. For small samples the interval is based on a t distribution, and 3.92 should be replaced by twice the t value. With 20 participants and an interval of 17.0 to 23.0, the multiplier is 2 x 2.093, and the standard deviation is 6.41.

Standard deviation from an interquartile range. If the data are roughly symmetric, the standard deviation is approximately the interquartile range divided by 1.35. An interquartile range of 10 gives about 7.4. This is an approximation that assumes normality, and studies where it is used should be examined in a sensitivity analysis. Better methods exist for medians with ranges or quartiles, and they are described in the literature on estimating means and standard deviations.

Combining two groups. A multi-arm trial may have to be combined into one group, for example two active arms against a single control. With means of 4.0 (SD 2.0, n = 30) and 6.0 (SD 3.0, n = 40), the combined mean is the weighted average, 5.14, and the combined standard deviation must include the variance between the group means: 2.79. Simply averaging the standard deviations would give 2.5, which is too small. The Cochrane Handbook gives the formula.

Change scores. The standard deviation of a change from baseline depends on the correlation between baseline and follow-up scores, which is rarely reported. Do not mix final values and change scores without thought, and consider a sensitivity analysis on the assumed correlation.

Multiple reports, arms, outcomes and time points

Multiple reports. Extract from all reports of a study together, recording each source, and use the most complete or the most recent data. Where reports conflict, note it and contact the authors.

Multiple arms. If two arms are eligible, either combine them or split the shared control group among comparisons, so that participants are not counted twice. Do not enter the control group in full for each comparison.

Multiple outcomes and scales. If several measures of the same outcome are reported, choose according to a rule written in the protocol, such as the primary outcome of the trial, the most commonly used instrument, or the one with the best validity. Choosing the most favorable result after looking at the data is a source of bias.

Multiple time points. Define in advance the time points of interest, for example short term (up to three months) and long term. Take the time point closest to the target when exact matches are not available, and record what was used.

Intention-to-treat and per-protocol. Prefer the analysis that includes all randomized participants, and record which was extracted. Record the number analysed, which often differs from the number randomized.

Where to find the data

Published articles are only one source. Trial registries often hold results tables, and regulatory documents can include far more detail than the journal article. Supplementary files, clinical study reports and protocol documents may contain outcome data. Figures can be digitized when numbers are absent, using software that reads values off graphs, with a record that the value was estimated from a figure and, preferably, extraction by two people. Contact the corresponding author when data are missing, unclear or inconsistent, give a clear list of what is needed and allow a reasonable time, and report the response rate.

Recording where each number came from

For every value, record the report, the page, table or figure, and whether it was reported directly, calculated or estimated. This lets another person verify it, makes later corrections straightforward and allows sensitivity analyses that exclude derived or estimated data. Keep the original completed forms and the final dataset with a version history. Many journals and funders now expect the extracted dataset and analysis code to be shared, and a clear audit trail makes this easy.

Common mistakes

  • Confusing a standard error with a standard deviation.
  • Extracting from only one report of a study with several.
  • Taking an unadjusted effect from one study and an adjusted one from another without stating this.
  • Entering the same control group several times.
  • Using the wrong direction of a scale, so that benefit appears as harm.
  • Choosing outcomes after seeing results.
  • Mixing time points without recording them.
  • Not recording the source of derived values.

Extraction for different kinds of review

The form changes with the question. For a review of prognostic factors, extract the adjustment variables used in each model, since adjusted estimates are only comparable when the sets of covariates are similar, and the CHARMS checklist lists the items. For a diagnostic accuracy review, extract the counts of true positives, false positives, false negatives and true negatives at each reported threshold, and the reference standard. For a network meta-analysis, extract the characteristics that may modify the treatment effect, such as baseline severity and prior treatment, for every arm, so that transitivity can be judged. For dose-response analysis, extract the dose or exposure level of each category with the number of cases and participants or person-years. For a qualitative synthesis, extraction centers on the context, the participants and the themes or findings, with illustrative quotations, and a different form is needed. For a scoping review, extraction is usually descriptive, mapping the characteristics of each study without numerical results.

Checking the extracted dataset

Once extraction is complete, run checks before analysis. Plot each effect and look for implausible values. Compare sample sizes with those in the paper. Verify that confidence intervals are consistent with the point estimates and sample sizes, since an interval that does not match is often a sign of an extraction or reporting error. Sort by effect size and check the extremes against the source. Compare the pooled result with that of earlier reviews on the same question, and investigate large differences. A random sample of studies can be re-extracted by a third person to estimate the error rate, which can then be reported. These checks find mistakes that would otherwise appear as unexplained heterogeneity or outliers. When an error is found, correct it in the dataset, note the change in the log and check whether the same mistake was made elsewhere, since errors tend to repeat within a paper or by a single extractor.

How we can help

We can design and pilot extraction forms, carry out independent double extraction, derive missing statistics with documented methods, contact authors, link multiple reports and deliver a clean, annotated dataset for analysis. [OWNER VERIFICATION REQUIRED] The relevant services are systematic review, meta-analysis and statistical analysis.

Frequently asked questions

Do I need two people to extract data?

Independent double extraction is the safest standard, at least for outcome data. A reduced approach with a second person checking is used in some rapid reviews and should be stated.

How do I get a standard deviation from a confidence interval?

Divide the interval width by 3.92 to get the standard error for a large sample, then multiply by the square root of the sample size. For small samples use the t distribution value in place of 1.96.

What if the study reports medians?

Approximate means and standard deviations with published methods, which assume roughly symmetric data, and test their effect in a sensitivity analysis.

Should I contact authors?

Yes, when data are missing, unclear or inconsistent. Be specific, allow time and report how many replied.

How do I handle a multi-arm trial?

Combine eligible arms, or split the shared control group, so that no participants are counted twice.

What should I record besides the value?

The report, page or table, and whether the value was reported, calculated or estimated from a figure.

References

  1. Li T, Higgins JPT, Deeks JJ. Chapter 5: Collecting data. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
  2. Buscemi N, Hartling L, Vandermeer B, Tjosvold L, Klassen TP. Single data extraction generated more errors than double data extraction in systematic reviews. J Clin Epidemiol. 2006;59(7):697-703.
  3. Mathes T, Klassen P, Pieper D. Frequency of data extraction errors and methods to increase data extraction quality: a methodological review. BMC Med Res Methodol. 2017;17:152.
  4. Wan X, Wang W, Liu J, Tong T. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Med Res Methodol. 2014;14:135.
  5. Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA. Chapter 10: Analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
  6. Higgins JPT, Eldridge S, Li T. Chapter 23: Including variants on randomized trials. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
  7. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.