Guide

What is meta-analysis?

A meta-analysis is a statistical method that combines the results of several independent studies addressing the same question, to produce one overall estimate and to examine why the results of the studies differ. This guide explains how it works, how to read the result, and where it can mislead, with a worked example.

Definition in plain terms

A meta-analysis is a statistical method for combining the quantitative results of separate studies that address the same question. Each study supplies an estimate of an effect, such as the difference in an outcome between a treatment group and a control group, together with a measure of how precisely that effect was estimated. The meta-analysis gives each estimate a weight, usually larger for more precise studies, averages them, and reports a single pooled estimate with a confidence interval.

Combining studies serves three purposes. It increases precision, because individual studies are often too small to estimate an effect reliably. It allows consistency to be checked, since an effect seen in one setting can be compared with the effect in others. And it makes disagreement between studies measurable and explorable instead of leaving it as an impression. The name reflects the idea of an analysis of analyses: the data being analyzed are the results of earlier studies and not the raw observations of individual participants, although meta-analysis of participant-level data is also possible.

A meta-analysis is a method of analysis, not a research design in itself. It is usually carried out as part of a systematic review, which is the structured process of finding, selecting and appraising the studies whose results are then combined. The distinction is explained in meta-analysis versus systematic review. A meta-analysis that pools whichever studies were convenient to find has no protection against selective inclusion, and its pooled estimate describes only those studies.

A short history

Pooling results across studies is older than the term. In 1904 the statistician Karl Pearson combined data from several studies of typhoid inoculation to judge its effect, which is often cited as an early example. In the 1930s Ronald Fisher described a method for combining the probabilities from independent tests, which dealt with significance and not with the size of effects.

The word meta-analysis was introduced by Gene Glass in 1976, and the approach was applied the following year by Smith and Glass to a large body of psychotherapy outcome studies. Through the 1980s and 1990s the methods spread through clinical medicine, supported by organizations such as Cochrane, which developed shared standards for searching, appraising and pooling trials. Today meta-analysis is used across the health sciences and in psychology, education, management, economics and ecology, and the statistical methods have developed considerably, including random-effects models, prediction intervals, network meta-analysis and Bayesian approaches.

How a meta-analysis works

The statistics are the last part of a longer process. A meta-analysis is only as sound as the steps that lead up to it.

  1. Define the question

    The question is stated precisely, including the population, the intervention or exposure, the comparator and the outcome. Eligibility criteria follow from it and are written down before searching begins.

  2. Find and select studies

    Databases and registries are searched with documented strategies, and studies are screened against the eligibility criteria. The process is reported with a flow diagram, as set out in PRISMA 2020.

  3. Extract effect estimates

    From each study, the effect estimate and its standard error are extracted, or computed from the reported counts, means and standard deviations. Study characteristics and risk-of-bias judgments are recorded alongside.

  4. Weight and pool

    Each study receives a weight that reflects the precision of its estimate, and the weighted average is computed under a fixed-effect or random-effects model. See fixed-effect and random-effects models.

  5. Examine differences and bias

    Heterogeneity between studies is quantified and explored, sensitivity analyses test the robustness of the result, and small-study effects are assessed. The certainty of the overall conclusion is then judged, for example with GRADE.

A worked example

The example below uses six simulated studies. The numbers were generated for illustration and do not come from real research. Each study reports a standardized mean difference (SMD), where a positive value favors the intervention, together with its standard error.

Simulated studies and their weights
StudyEffect (SMD)Standard errorFixed-effect weightRandom-effects weight
Study A0.300.1513.9%14.8%
Study B0.100.1221.6%21.2%
Study C0.450.207.8%9.0%
Study D0.220.1031.2%27.6%
Study E-0.050.189.6%10.9%
Study F0.380.1415.9%16.6%

In a fixed-effect analysis each study is weighted by the inverse of its variance, so Study D, with the smallest standard error, carries the most weight at 31.2 percent and Study C, with the largest, carries the least at 7.8 percent. The weighted average is an SMD of 0.22, with a 95 percent confidence interval from 0.11 to 0.33.

A random-effects analysis first estimates how much the true effects vary between studies. Here the Q statistic is 6.16 on 5 degrees of freedom, I² is 19 percent and the between-study variance (τ²) is 0.0045. Because the true effects are allowed to differ, each study's weight is adjusted so that the weights become more even, as the last column shows. The pooled SMD is 0.22 with a 95 percent confidence interval from 0.10 to 0.35, which is wider than the fixed-effect interval because it carries the uncertainty about between-study variation.

Forest plot of the six simulated studies in the worked exampleSix studies are shown as squares with horizontal lines for their 95% confidence intervals. The random-effects pooled estimate is a diamond at 0.22, with a 95% confidence interval from 0.10 to 0.35. The data are simulated.StudySMD [95% CI]Study A0.30 [0.01, 0.59]Study B0.10 [-0.14, 0.34]Study C0.45 [0.06, 0.84]Study D0.22 [0.02, 0.42]Study E-0.05 [-0.40, 0.30]Study F0.38 [0.11, 0.65]Random effects0.22 [0.10, 0.35]-0.50.00.51.0Favors controlFavors intervention
Forest plot of the worked example. Each square is a study, with a size that reflects its random-effects weight, and each line is its 95% confidence interval. The diamond is the pooled estimate. The data are simulated.

The result has a second interpretation that is easy to miss. The 95 percent prediction interval, which describes where the true effect in a new study would be expected to fall, runs from -0.03 to 0.48 and includes zero. The pooled average is clearly above zero, but a future study in a different setting could still show no benefit or a small harm. A reader who looked only at the confidence interval would reach a more confident conclusion than the data support.

How to read a meta-analysis

Pooled estimate
The weighted average effect across studies. Its meaning depends on the effect measure: a ratio of 1 or a difference of 0 means no effect.
Confidence interval
The range of values for the average effect that are compatible with the data. A narrow interval signals precision, not correctness.
Heterogeneity
How much the true effects differ between studies, usually summarized by I² and the between-study variance. See heterogeneity and I².
Prediction interval
The range in which the true effect of a new study would be expected to fall. See prediction intervals.
Forest plot
The standard display: one line per study, with a diamond for the pooled result. See how to read a forest plot.
Funnel plot
A graph of each study's effect against its precision that can reveal small-study effects, including publication bias. See funnel plots.

Beyond these outputs, the reader should check three things that the numbers do not show: how the studies were found, how their risk of bias was assessed, and how many of the included studies contribute to each conclusion. A tidy forest plot built on a biased set of studies is still biased.

Types of meta-analysis

The basic procedure has many extensions, each suited to a different kind of question or data.

How meta-analysis is used across fields

The questions differ by discipline, although the logic is the same. The examples below are illustrative questions, not findings.

  • Medicine and health sciences: does a drug reduce mortality compared with a placebo, or how accurate is a screening test?
  • Psychology: how large is the average effect of a therapy on symptoms, and does it depend on the population?
  • Education: what is the average effect of a teaching method on achievement, and does it vary by grade level?
  • Management: how strongly are two organizational constructs related across firms and settings?
  • Economics: how responsive is labor supply to changes in wages, as estimated across studies?
  • Ecology and life sciences: how does an environmental change affect growth or survival across experiments and species?

Common misunderstandings

More studies always give a better answer
Adding studies that share the same bias gives a more precise estimate of a biased quantity. The quality and independence of the studies matter as much as their number.
A meta-analysis is automatically the strongest evidence
It is only as strong as the studies it combines and the process that selected them. Certainty ratings such as GRADE exist because pooled evidence can still be of low certainty.
It combines apples and oranges
This is a fair concern when studies ask different questions. It is addressed by precise eligibility criteria and by examining heterogeneity, and in some cases the right decision is not to pool.
A non-significant pooled result means there is no effect
It may reflect an imprecise estimate. The confidence interval shows which effect sizes remain plausible, including clinically or practically important ones.
Counting significant studies is an adequate summary
Vote counting ignores the size and precision of each study and can mislead in either direction. Weighted pooling does not have this flaw.
Heterogeneity is simply a defect
Differences between studies are information. They can show where an effect is larger or smaller and prompt better questions.

Strengths and limitations

The main strengths are precision, transparency and the ability to examine consistency across settings. A well-conducted meta-analysis shows how a conclusion depends on the choices behind it, and it makes disagreement between studies visible.

The main limitations follow from its dependence on the underlying studies. Publication bias and selective reporting can distort the available evidence. Bias that affects all included studies in the same way cannot be corrected statistically. Observational studies are exposed to confounding that no pooling can remove. With few studies, estimates of between-study variation are imprecise and several diagnostic checks lose power. A pooled number can also hide meaningful differences between subgroups. These are reasons to report a meta-analysis carefully, and sometimes reasons to choose a different synthesis method; see methods for the alternatives.

Planning your own meta-analysis

A sound project begins with a protocol that fixes the question, eligibility criteria, outcomes and analysis plan before the search is run, and registers it where a suitable registry exists. The next stages are the search, screening, data extraction, risk-of-bias assessment, analysis and reporting. Each stage has its own guidance, listed under the resources, and the reporting is checked against PRISMA 2020 and, for observational data, MOOSE.

The meta-analysis service and the systematic review service describe the support available at each stage.

Get a quoteDescribe your question, study type and target journal.

Frequently asked questions

What is the difference between a meta-analysis and a systematic review?

A systematic review is the structured process of identifying, appraising and summarizing all relevant studies on a question. A meta-analysis is the statistical technique for combining their numerical results. Many systematic reviews include a meta-analysis, but a review can be reported without one when the studies are too different to pool.

How many studies are needed for a meta-analysis?

There is no fixed minimum. Two studies can be combined mathematically, but with few studies the estimate of between-study variation is imprecise and tests for small-study effects are unreliable. Whether pooling is meaningful depends on how comparable the studies are as much as on how many there are.

Is a meta-analysis the same as a literature review?

No. A literature review summarizes existing work, often selectively and without a defined method. A meta-analysis is a statistical combination of results, and when it is part of a systematic review it follows a documented search and explicit eligibility criteria.

Can a meta-analysis combine observational studies?

Yes, usually by pooling adjusted effect estimates and reporting against the MOOSE guideline. The main concern is confounding, which pooling does not remove, so the interpretation is more cautious than for randomized trials.

What software is used to run a meta-analysis?

Common choices include R, with packages such as metafor and meta, Stata, RevMan and Comprehensive Meta-Analysis. The choice rarely changes the result when the same methods are used, but the software and settings should be reported.

Is a meta-analysis always the best form of evidence?

No. It depends on the quality of the included studies and the way they were selected. A large, well-conducted trial can be more reliable than a meta-analysis of small, biased ones, which is why the certainty of the evidence is assessed separately.

References

  1. Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  2. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
  3. Cooper H, Hedges LV, Valentine JC, editors. The Handbook of Research Synthesis and Meta-Analysis. 3rd ed. Russell Sage Foundation; 2019.
  4. Glass GV. Primary, secondary, and meta-analysis of research. Educ Res. 1976;5(10):3-8.
  5. Smith ML, Glass GV. Meta-analysis of psychotherapy outcome studies. Am Psychol. 1977;32(9):752-760.
  6. Pearson K. Report on certain enteric fever inoculation statistics. BMJ. 1904;2:1243-1246.
  7. Fisher RA. Statistical Methods for Research Workers. 4th ed. Oliver and Boyd; 1932.
  8. DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials. 1986;7(3):177-188.
  9. Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med. 2002;21(11):1539-1558.
  10. Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.
  11. IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6(7):e010247.
  12. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
  13. Stroup DF, Berlin JA, Morton SC, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-2012.
  14. Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924-926.
  15. Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1-48.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.