Method

Meta-analysis: the method

Meta-analysis is the statistical combination of results from separate studies. This page describes the method itself: the data it needs, the models used, the quantities it reports, the assumptions behind it, and the decisions an analyst must make. For a plain introduction, see the guide to what meta-analysis is.

What the method covers

Meta-analysis is a family of statistical procedures that share one idea: results from several studies of the same question are combined into a summary that is more informative than any single study. The common core is a weighted average of study-level effect estimates. Everything else, including the choice of effect measure, the model for between-study differences, the estimation method, the handling of dependent data and the checks on bias, varies with the question and the data. The method is used across medicine, psychology, education, management, economics and ecology, with different conventions in each.

This page is a reference to the method. It complements the introductory guide, which explains the idea with a worked example, and the meta-analysis service, which describes how analyses are carried out for clients. The specialized forms of the method each have a page of their own, linked in the sections below.

Data and effect measures

A meta-analysis works on study-level summaries. For each study the analyst needs an effect estimate and a variance or standard error, on a scale where averaging is meaningful. The effect measure should suit the outcome. For binary outcomes the usual measures are the risk ratio, the odds ratio and the risk difference. For continuous outcomes they are the mean difference, when all studies use the same scale, and the standardized mean difference, when they use different scales for the same construct. For time-to-event outcomes it is the hazard ratio. For associations it is the correlation coefficient, usually pooled after Fisher's z transformation, and for single-group data the proportion or the rate.

Ratio measures are analyzed on the log scale, because the log of a ratio is approximately symmetric and normally distributed and its standard error is straightforward. Results are transformed back for presentation. The choice among measures involves interpretability and statistical properties. Odds ratios are convenient for modeling and are the natural output of case-control studies, but they overstate risk ratios when the outcome is common and are less intuitive. Standardized mean differences make scales comparable at the price of results that depend on the variability of each sample. The guide on effect sizes compares the measures.

Weighting and the pooled estimate

The pooled estimate is a weighted average of the study estimates. In the inverse-variance method each study receives a weight equal to the reciprocal of its variance, so that more precise studies count for more. If study i has estimate y_i and variance v_i, the weight is w_i = 1/v_i, the pooled estimate is the sum of w_i times y_i divided by the sum of w_i, and its variance is 1 divided by the sum of w_i. A 95 percent confidence interval is the estimate plus and minus 1.96 times its standard error. This is the fixed-effect, or common-effect, analysis, described on its own page as fixed-effect meta-analysis.

Other weighting schemes exist for particular data. The Mantel-Haenszel method pools binary outcomes using weights computed directly from the counts, and it behaves well when data are sparse. The Peto method pools odds ratios from observed and expected events and suits rare events with balanced groups. For correlations and for proportions, transformations are applied before the inverse-variance step. In every case the principle is the same: studies are weighted according to the information they contain.

Fixed-effect and random-effects models

The central modeling decision is how to treat differences between studies. A fixed-effect model assumes that all studies estimate one common true effect, so observed differences reflect sampling error only. A random-effects model assumes that the true effects vary between studies and are drawn from a distribution, and it estimates the mean of that distribution and its variance, written tau squared. Under random effects each study's weight becomes 1 divided by the sum of its own variance and tau squared, which makes the weights more equal and widens the confidence interval.

The two models answer different questions. The fixed-effect result describes the studies actually included, on the assumption that they are estimating the same thing. The random-effects result describes the average of a population of effects that the included studies sample, and is the more defensible basis for generalizing when the studies differ in populations, settings or methods, which they nearly always do. The choice should be reasoned, and where it matters the other model is shown as a sensitivity analysis. See random-effects meta-analysis and fixed-effect and random-effects models.

Heterogeneity

Heterogeneity is variation in true effects between studies beyond chance. It is assessed with the Q statistic, the sum of the weighted squared deviations of the study estimates from the pooled estimate, which is compared with a chi-squared distribution with k minus 1 degrees of freedom for k studies. The between-study variance, tau squared, quantifies it on the scale of the effect measure, and I squared expresses the proportion of total variability attributable to heterogeneity as a percentage. These quantities have limits. The Q test has low power with few studies and excessive power with many. I squared depends on the precision of the studies and rises toward 100 percent as studies get larger, even when the differences are small in practice. Tau squared is imprecisely estimated when studies are few.

A prediction interval, which gives the range in which the true effect of a new study would be expected to fall, is more informative than any of these for many purposes. When heterogeneity is present, it should be explored, with subgroup analysis or meta-regression, and interpreted, not simply reported. See the guides on heterogeneity, I squared and prediction intervals.

Steps in carrying out an analysis

  1. Specify the analysis in advance

    The effect measure, the model, the handling of missing and dependent data, and the planned subgroup and sensitivity analyses are written into the protocol before the results are seen.

  2. Prepare the data

    Effect estimates and variances are computed or extracted, conversions are documented, multi-arm and clustered designs are handled correctly, and the dataset is checked.

  3. Fit the main model

    The pooled estimate, its interval and the heterogeneity statistics are computed with the specified model and estimator.

  4. Examine differences

    Subgroup analysis or meta-regression is carried out where the plan calls for it, and a prediction interval is reported.

  5. Test robustness

    Sensitivity analyses vary the model, the estimator, the effect measure and the set of studies, and influence diagnostics identify studies that drive the result.

  6. Assess bias

    Small-study effects and publication bias are assessed when enough studies are available, and the certainty of the conclusion is rated.

Assumptions

Meta-analysis rests on assumptions that should be examined and stated. The studies are assumed to be independent, which fails when they share participants, authors or a control group, or when one study contributes several estimates. The effect estimates are assumed to be approximately normally distributed with known variances, which is a good approximation for large studies and poor for small ones or for sparse binary data. Under random effects, the true effects are assumed to be normally distributed, which is hard to check with few studies. The studies are assumed to be comparable enough that an average is meaningful, which is a matter of judgment about populations, interventions and outcomes and not of statistics. And the set of studies is assumed to be representative of the available evidence, which publication bias and selective reporting can violate.

Violations do not always invalidate an analysis, but they change how it should be done and read. Dependent estimates call for multilevel or robust variance models. Sparse data call for exact methods. A set of studies that is not comparable may call for no pooling at all.

Bias and sensitivity

A pooled estimate inherits the biases of its studies. Risk of bias within studies is assessed with design-specific tools and used in sensitivity analyses. Bias across studies arises when the available evidence is a selected sample of all that was done. A funnel plot of the effect estimates against their precision, and tests of its asymmetry, can signal small-study effects, which include publication bias but also real differences between small and large studies, and are recommended only when about ten or more studies are available. Methods such as trim-and-fill and selection models are used as sensitivity analyses. See publication bias and sensitivity analysis.

Forms of the method

The basic procedure has many extensions, each designed for a kind of question or data.

Software

Meta-analysis can be carried out in general statistical software and in dedicated tools. In R, the metafor and meta packages provide the full range of models, and other packages cover specialized forms. Stata has a meta-analysis suite and user-written commands. RevMan, the software of Cochrane, supports standard reviews of interventions, and other programs, including commercial ones, offer point-and-click analysis. The results of the same method should not depend on the software, but the defaults differ, for instance in the estimator of tau squared or the confidence interval method, so the settings should be reported and not left at their defaults without comment. Analyses are best done in scripts that can be re-run.

Reporting

The report should let a reader judge and, in principle, repeat the analysis. It states the effect measure and the model, the estimator of between-study variance and the confidence interval method, the software and its version, the heterogeneity statistics and the prediction interval, the planned subgroup and sensitivity analyses, and any analyses that were not planned. The pooled results are presented in a forest plot with the individual studies, with the number of studies and participants. The review is reported against PRISMA 2020, and meta-analyses of observational studies are also reported against MOOSE. See PRISMA 2020.

Common pitfalls

  • Pooling studies that ask different questions, so that the average describes nothing in particular.
  • Treating dependent estimates as independent, which gives intervals that are too narrow.
  • Choosing the model, the estimator or the effect measure after seeing which gives the preferred result.
  • Reading I squared as a fixed verdict on whether pooling is appropriate.
  • Reporting a significant pooled result without the prediction interval or the certainty of the evidence.
  • Interpreting subgroup or meta-regression findings from a few studies as established.
  • Ignoring publication bias, or testing for it with too few studies to be meaningful.
  • Treating a pooled number as proof, when it is only as good as the studies behind it.

How we can help

Support

Support for meta-analysis

The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.

Get a quoteTell us your question, the number of studies and your target journal.

Frequently asked questions

What is the difference between meta-analysis and pooling data?

A meta-analysis combines study-level results, usually effect estimates and variances, with weights that reflect their precision, and it keeps track of which study each result came from. Simple pooling of raw counts across studies can give misleading results because it ignores differences between studies.

Why are effect sizes weighted by their variance?

More precise estimates carry more information, so giving them more weight yields the most precise combined estimate. Under random effects the weights also include the between-study variance, which makes them more equal.

Do I need to use the log scale for ratios?

Ratio measures are analyzed on the log scale because the sampling distribution is closer to normal and the variance is simpler. Results are transformed back for reporting.

Is I-squared enough to decide whether to pool studies?

No. It depends on the precision of the studies and is imprecise with few studies. Judgment about whether the studies ask comparable questions is more important, and a prediction interval is usually more informative.

Can meta-analysis be done in Excel or point-and-click software?

It can, for simple analyses, but scripted analysis in R or Stata is better for transparency and reproducibility, and defaults differ between programs, so the settings must be reported.

How is a meta-analysis reported?

Against PRISMA 2020, with MOOSE added for observational studies, giving the model, estimator, software, heterogeneity, sensitivity analyses and a forest plot.

References

  1. Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  2. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
  3. Hedges LV, Olkin I. Statistical Methods for Meta-Analysis. Academic Press; 1985.
  4. Whitehead A. Meta-Analysis of Controlled Clinical Trials. Wiley; 2002.
  5. Cooper H, Hedges LV, Valentine JC, editors. The Handbook of Research Synthesis and Meta-Analysis. 3rd ed. Russell Sage Foundation; 2019.
  6. DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials. 1986;7(3):177-188.
  7. Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med. 2002;21(11):1539-1558.
  8. Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.
  9. Rucker G, Schwarzer G, Carpenter JR, Schumacher M. Undue reliance on I-squared in assessing heterogeneity may mislead. BMC Med Res Methodol. 2008;8:79.
  10. Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1-48.
  11. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.