Data extraction and coding

We extract and code the data that a meta-analysis depends on: designing the extraction form, taking results from the studies, computing effect sizes with documented conversions, and checking the dataset. Errors at this stage carry through to the pooled estimate, and they are among the most common reasons for wrong results.

Why extraction matters

Extraction is the step in which the numbers in published papers become the data of a review. It looks routine and is anything but. Papers report results in different ways, with different units, time points and effect measures, and often with errors or omissions. The extractor must decide which of several reported results to use, convert it to a common form, and enter it correctly. Each of these decisions can change the pooled estimate.

The error rate in extraction is not small. Studies of published meta-analyses have found extraction errors in a substantial share of them, including cases in which correcting the errors changed the conclusion, and meta-research has shown that independent double extraction reduces errors compared with a single extractor. For this reason extraction is treated as a measurement process that needs a protocol, a form, a check and a record. It is part of the systematic review process and is the input for the meta-analysis.

What the service includes

  • Extraction form and codebook. A form designed for the review, with definitions of every data item, piloted on a sample of studies.
  • Extraction of study characteristics, participants, interventions or exposures, outcomes and results.
  • Independent verification. Double extraction, or independent checking of key data, as set in the protocol.
  • Effect-size computation, with each conversion documented.
  • Linking of reports to studies, so that each study is counted once.
  • Queries to authors where data are missing and the request is proportionate.
  • Dataset audit for entry errors, implausible values and inconsistencies.
  • Delivery of a clean analysis-ready dataset, with a log of every decision.

Designing the extraction form

A good form is specific to the review, as short as it can be and unambiguous. Every item has a definition, and the allowed values are listed, so that two extractors would enter the same thing. The form usually captures the identity of the study and its reports, the design, setting and dates, the characteristics of the participants, the details of the intervention or exposure and the comparator, the definitions and timing of the outcomes, the results needed for the analysis, the funding and declared conflicts, and the risk-of-bias information if the appraisal is done on the same form. Items that will not be used in the analysis or the report are left out, because each one adds time and error.

The form is piloted on a few studies of different kinds by the people who will use it. Problems that appear, such as unclear definitions or results that fit nowhere, are fixed before the main extraction. The final version is fixed and any later change is recorded, with the studies already extracted updated to match.

Extraction in practice

Two people extract independently from the full texts and the results are compared, with discrepancies resolved by discussion or a third person. Where resources are limited, a documented alternative, such as one extractor with a second person checking all the data used in the analysis, can be used, and its risk is acknowledged. The data should be taken from the most complete source: the main paper first, then supplements, registry entries and any additional reports, with a rule for deciding which to use when they disagree.

The unit of extraction is the study, not the paper. One trial may be reported in several papers, and one paper may report several studies. Linking reports to the parent study avoids counting the same participants twice, which would inflate the weight of that study. Where it is unclear whether two reports share participants, the authors are asked, and the uncertainty is handled in a sensitivity analysis.

Computing effect sizes

The analysis needs, for each study and outcome, an effect estimate and a measure of its precision. Studies report a bewildering variety of formats, and the conversion between them is part of the work. For binary outcomes the counts of events and participants in each group give a risk ratio, odds ratio or risk difference. For continuous outcomes the means, standard deviations and sample sizes give a mean difference or a standardized mean difference, with a correction for small samples for Hedges' g. Correlations are transformed to Fisher's z for pooling. Hazard ratios and their standard errors are taken from the paper or derived from confidence intervals. See the guide on effect sizes.

Formulas exist for many conversions: a standard error from a confidence interval, a standard deviation from a standard error, and effect sizes from t statistics or p-values when nothing else is reported. They are applied with care, because each depends on assumptions, and they are recorded so that the reader can see which values were reported and which were derived. A sensitivity analysis that excludes studies that required substantial conversion shows how much they matter.

Awkward cases

Medians and ranges
Studies often report a median with a range or quartiles instead of a mean and standard deviation. Methods exist for estimating the mean and standard deviation from these summaries, and they work better for some distributions and sample sizes than for others, so they are used with caution and flagged.
Multi-arm trials
A trial with several intervention arms and one control cannot enter an analysis as several independent comparisons without double-counting the control group. The arms are combined or the control group is split, according to the question.
Change from baseline and final values
Studies report one or the other, and a pooled analysis must not mix them without care, because their variances differ. The choice and the handling of the correlation between baseline and final values are documented.
Cluster and crossover designs
Analyses that ignore clustering or pairing give standard errors that are wrong, and adjustments, such as the design effect for clusters, need information the paper may not give.
Survival data
When a hazard ratio is not reported, it may be estimated from reported statistics or reconstructed from published curves. These routes are less reliable and should be flagged.
Data only in figures
Values can be read from graphs with digitizing software, which adds measurement error that must be recorded.

Coding study characteristics

Beyond the outcome data, a review often codes characteristics of the studies for use in subgroup analysis, meta-regression and description: the population, the setting, the dose or duration of the intervention, the year, the country, the design features and the risk of bias. Coding turns descriptions into variables, and the way it is done affects what can be analyzed. Categories are defined in advance with explicit rules and examples, continuous characteristics are recorded as numbers where the paper allows and not forced into bands, and the codes for rare categories are decided with the analysis in mind, since a category with one or two studies cannot be analyzed. Where coding involves judgment, two coders work independently on a sample, and their agreement is measured, typically with a chance-corrected statistic, and disagreements are resolved and the rules clarified. The coded variables feed the analyses described under meta-regression.

Data management and version control

The dataset is a research record, and it is managed as one. Raw extraction entries are kept separate from the cleaned analysis dataset, so that every change can be traced. Each version of the dataset is labeled with a date and a description, and the analysis code is tied to a specific version. Study and report identifiers are consistent across files, and the codebook is kept with the data so that the meaning of every variable is recorded. The data are stored securely, and where they include any information that could identify people, for example in individual participant data projects, they are handled according to the permissions and the privacy rules that apply. These habits are unglamorous, and they are what makes a result reproducible months later, when a reviewer asks a question about a single study.

Checking the dataset

Before the data go to the analysis, they are audited. Checks include verification of a sample, or all, of the entered values against the source; range and plausibility checks, such as standard deviations that are implausibly small; consistency between related values, such as the sum of group sizes and the total reported; comparison of the extracted effect with any effect reported by the authors; and a check for duplicate studies. Discrepancies are traced to the source and resolved, and the log records the change and the reason. The dataset is then frozen and version-labeled, so that the analysis can be tied to a specific version.

Deliverables

  • Extraction form and codebook with definitions and allowed values.
  • Extraction dataset, in a format ready for analysis, with study and report identifiers.
  • Effect sizes and standard errors, with a record of each conversion and its basis.
  • Discrepancy and query log, and record of contact with authors.
  • Audit report on the checks made and their results.

Get a quoteTell us the number of studies and the outcomes.

Optional extension: from dataset to published paper

Optional extension

Carry the dataset through to a full analysis and paper

When the scope continues beyond the dataset, the package also contains the following.

Get a quoteTell us your scope, target journal and deadline.

Manuscript preparation follows the research integrity and authorship statement. The researchers who conceived the study and interpret its findings remain responsible for the content and its conclusions, and contributions that do not meet authorship criteria are acknowledged.

Limitations

Extraction cannot improve the reporting of the studies. When a paper does not give the data, they cannot be invented, and where an estimate has to be derived, its uncertainty is passed to the analysis. Different reasonable extractors can make different choices when a paper reports several results for an outcome, so rules should be set in advance and the choices recorded. Contact with authors helps, but replies are incomplete and slow. Because extracted data are the foundation of the analysis, errors cannot be eliminated, only reduced and documented, which is why the audit and the sensitivity analyses matter. The service does not select results to favor a conclusion, and every conversion is recorded.

This service provides research and evidence-synthesis support. It does not provide clinical advice.

How long does it take?

The time depends on the number of studies and outcomes, how clearly the papers report their results, how many conversions are needed, whether double extraction is used, and how many author queries are sent. Studies that report results in full are quick to extract, and those that report little take far longer. After a look at a sample of the studies, a schedule and its dependencies are agreed.

Frequently asked questions

Is double extraction necessary?

It is recommended because extraction errors are common and can change conclusions. Where resources are limited, a documented alternative, such as a second person checking all the data used in the analysis, can be used, with the limitation stated.

What do you do when a study reports a median instead of a mean?

Methods exist for estimating the mean and standard deviation from medians and ranges or quartiles. They rely on assumptions, so the conversion is flagged, and a sensitivity analysis shows whether such studies affect the result.

How do you avoid counting a study twice?

Reports are linked to their parent study during extraction, using trial registration numbers, author lists and participant details. Where overlap is unclear, the authors are asked and the uncertainty is explored in a sensitivity analysis.

Do you contact study authors?

Where missing data matter and a request is proportionate, authors are contacted with specific questions. Replies are recorded, and the lack of a reply is not treated as information.

Can you extract data from figures?

Values can be read from graphs with digitizing software, which adds measurement error that is recorded and, if needed, examined in a sensitivity analysis.

Will you use only the results that support my hypothesis?

No. Results are extracted according to rules set in the protocol, and the choice among several reported results is made in advance.

References

  1. Li T, Higgins JPT, Deeks JJ. Chapter 5: Collecting data. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  2. Gotzsche PC, Hrobjartsson A, Maric K, Tendal B. Data extraction errors in meta-analyses that use standardized mean differences. JAMA. 2007;298(4):430-437.
  3. Mathes T, Klassen P, Pieper D. Frequency of data extraction errors and methods to increase data extraction quality: a methodological review. BMC Med Res Methodol. 2017;17:152.
  4. Jones AP, Remmington T, Williamson PR, Ashby D, Smyth RL. High prevalence but low impact of data extraction and reporting errors were found in Cochrane systematic reviews. J Clin Epidemiol. 2005;58(7):741-742.
  5. Wan X, Wang W, Liu J, Tong T. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Med Res Methodol. 2014;14:135.
  6. Luo D, Wan X, Liu J, Tong T. Optimally estimating the sample mean from the sample size, median, mid-range, and/or mid-quartile range. Stat Methods Med Res. 2018;27(6):1785-1805.
  7. Shi J, Luo D, Weng H, et al. Optimally estimating the sample standard deviation from the five-number summary. Res Synth Methods. 2020;11(5):641-654.
  8. Hozo SP, Djulbegovic B, Hozo I. Estimating the mean and variance from the median, range, and the size of a sample. BMC Med Res Methodol. 2005;5:13.
  9. Parmar MKB, Torri V, Stewart L. Extracting summary statistics to perform meta-analyses of the published literature for survival endpoints. Stat Med. 1998;17(24):2815-2834.
  10. Tierney JF, Stewart LA, Ghersi D, Burdett S, Sydes MR. Practical methods for incorporating summary time-to-event data into meta-analysis. Trials. 2007;8:16.
  11. Guyot P, Ades AE, Ouwens MJNM, Welton NJ. Enhanced secondary analysis of survival data: reconstructing the data from published Kaplan-Meier survival curves. BMC Med Res Methodol. 2012;12:9.
  12. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.