Method

Pairwise meta-analysis

Pairwise meta-analysis combines studies that compare the same two conditions, such as a treatment and a control. It is the most common form of the method and the foundation of network meta-analysis. This page covers the data it needs, the models and measures used, and the problems that arise in practice.

What pairwise meta-analysis is

A pairwise meta-analysis answers a question of the form: how does condition A compare with condition B, summarized across the studies that compared them. Each study provides a contrast between the two, for example the difference in mean blood pressure, the ratio of the risk of an event, or the hazard ratio for survival, together with its standard error. The analysis combines these contrasts with weights and reports a single pooled contrast with a confidence interval, an assessment of how much the studies differ, and the checks described below.

The term is used to distinguish this form from network meta-analysis, in which several interventions are compared through a connected set of trials. Pairwise analysis is the building block of that approach and the standard method in most published reviews of interventions. It is also used for non-intervention comparisons, such as exposed against unexposed groups in observational studies, or men against women in studies of sex differences, provided the contrast is defined in the same way in every study. The general approach is described on the page on meta-analysis.

When it is appropriate

Pairwise meta-analysis is appropriate when a set of studies compares the same two conditions on the same outcome and when the studies are similar enough that an average is meaningful. Three checks matter. The comparison must be the same: a trial of a drug against placebo and a trial of the same drug against an active comparator are different contrasts, and combining them as if they were one gives a result that answers neither question. The outcome must be defined and measured in a compatible way and at comparable time points. And the populations and settings must be similar enough that the effect is not expected to differ greatly, or that any differences can be examined.

When these conditions fail, other approaches are better. If several treatments are of interest, network meta-analysis uses all the trials in one model. If the studies are too different, subgroup analysis can split them into more homogeneous sets, or a narrative synthesis can be reported. And if only one study exists, a meta-analysis is not possible.

Data layout

The data for a pairwise meta-analysis can be arranged in two ways. In arm-level format, each study contributes a row for each group, with the counts, or the mean, standard deviation and sample size, in that group. In contrast-level format, each study contributes a single row with an effect estimate and standard error for the comparison. Arm-level data are richer and allow some analyses, such as exact models for binary outcomes and the choice of effect measure at the analysis stage, that contrast-level data do not. Contrast-level data are all that many publications provide, particularly for adjusted estimates and hazard ratios, and they are what the generic inverse-variance method uses.

For each study the dataset also records the characteristics needed for subgroup analysis and for appraisal: the population, the intervention and comparator, the follow-up, and the risk of bias. The layout and variable definitions are fixed in a codebook, and the checks on the data are described under data extraction and coding.

Choosing the effect measure

Binary outcome
Risk ratio, odds ratio or risk difference. The risk ratio is easiest to interpret, the odds ratio has convenient modeling properties, and the risk difference is directly relevant to absolute benefit but varies more with baseline risk and is less likely to be consistent across studies.
Continuous outcome, same scale
Mean difference. It keeps the original units and is directly interpretable.
Continuous outcome, different scales
Standardized mean difference, with a small-sample correction in the form of Hedges' g. It removes the units but depends on the variability of each sample.
Time to event
Hazard ratio on the log scale, with its standard error from the paper or derived from other statistics.
Rates
Rate ratio, with person-time as the denominator.

The choice is made in the protocol, with the reason, and alternatives are examined as sensitivity analyses. Where baseline risk differs widely among studies, relative measures are generally more stable across studies than absolute ones, and absolute effects for a target baseline risk can be computed afterwards from the pooled relative effect. See the guides on odds and risk ratios and standardized mean differences.

Models

The pooled contrast is estimated with a fixed-effect or a random-effects model, as described in fixed-effect meta-analysis and random-effects meta-analysis. For binary outcomes with arm-level data, three families of methods are in routine use. The inverse-variance method works on log ratios and is the general-purpose option. The Mantel-Haenszel method uses the counts directly and has better properties when events are few or study sizes are small. The Peto method is based on observed minus expected events and suits rare events and small effects when group sizes are balanced, but is biased when there are large imbalances or large effects. More advanced approaches fit generalized linear mixed models, such as a binomial model with random effects, to the arm-level counts and avoid the need for continuity corrections.

The random-effects estimator of between-study variance, and the method of computing the confidence interval, matter when studies are few. Restricted maximum likelihood and the Paule-Mandel estimator are generally preferred to the older DerSimonian-Laird estimator in methodological comparisons, and the Hartung-Knapp-Sidik-Jonkman adjustment improves the coverage of the interval with few studies.

Multi-arm and non-standard trials

Real trials do not always fit the two-group template. A trial with several intervention arms and one control cannot enter as separate comparisons without counting the control group more than once, which gives that trial too much weight and ignores the correlation between its comparisons. The options are to combine the intervention arms that are relevant to the question into one, when this makes sense, or to split the control group among the comparisons, or to use a model that accounts for the correlation, as network meta-analysis does. The choice depends on whether the arms are alternative versions of the same intervention or distinct interventions.

Cluster-randomized trials, in which groups of people are randomized together, must be analyzed with allowance for clustering, because ignoring it gives standard errors that are too small. If the trial's own analysis did not account for it, the standard error can be inflated by the design effect, which needs an estimate of the intracluster correlation. Crossover trials require paired analysis, and results reported as if from parallel groups overstate the variance. Trials reporting several time points or several outcomes contribute dependent estimates and need a rule for choosing one, or a model for dependence. These issues are handled in the analysis plan.

Sparse data and zero events

When events are rare, studies may have zero events in one or both arms, which makes the usual log ratio undefined. The common remedy of adding a half to each cell is simple but can bias results, particularly when group sizes differ. Alternatives that avoid the problem include Mantel-Haenszel methods without continuity correction, the Peto method under its conditions, exact methods and random-effects models fitted directly to the counts. Studies with zero events in both arms contribute no information to a ratio measure in most methods and are sometimes dropped, which can matter when they form a large part of the evidence. Whatever is chosen, a sensitivity analysis with a different approach shows whether the conclusion depends on it. See the guide on zero-event studies.

Heterogeneity and exploring it

After the pooled contrast is estimated, the differences between studies are examined. The Q statistic, tau squared and I squared quantify the variation, and a prediction interval describes the range of effects to expect in a new study. If the differences matter, they are explored with subgroup analysis, which compares pooled contrasts across categories of a characteristic, and with meta-regression, which relates the contrast to the characteristic. A formal test for interaction, not a comparison of significance levels, is the basis for claiming a difference between subgroups. The pitfalls of such analyses, particularly with few studies, are covered in the page on meta-regression.

Bias and sensitivity

Sensitivity analyses test whether the conclusion survives reasonable changes: excluding studies at high risk of bias, using a different model or estimator, using a different effect measure, leaving out each study in turn, and excluding studies whose data required substantial derivation. Small-study effects are assessed with funnel plots and, when about ten or more studies are available, with tests suited to the effect measure. The certainty of the conclusion for each outcome is then rated, usually with GRADE. The results should be consistent with the risk-of-bias assessment: a pooled estimate based mainly on studies at high risk of bias deserves cautious wording.

Reporting

A pairwise meta-analysis is reported with a forest plot that shows each study's estimate and interval, its weight, and the pooled estimate, together with the number of studies and participants, the model and estimator, the heterogeneity statistics and the prediction interval. Subgroup and sensitivity results are presented in the text or tables, and any analysis not in the protocol is labeled as post hoc. The review follows PRISMA 2020, and the data and code are shared where possible. See how to read a forest plot.

How we can help

Support

Support for pairwise meta-analysis

The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.

Get a quoteTell us your question, the number of studies and your target journal.

Frequently asked questions

What is the difference between pairwise and network meta-analysis?

Pairwise meta-analysis compares two conditions using studies that compared them directly. Network meta-analysis compares several conditions at once, using direct and indirect evidence through a connected network of studies.

Can I combine trials that used different comparators?

Usually not in one pool, since a drug against placebo and the same drug against another active treatment are different comparisons. They can be analyzed separately, or in a network meta-analysis if the assumptions hold.

What do I do with a trial that has three arms?

Combine the arms that are alternative versions of the same intervention, or split the shared control group, or use a model that accounts for the correlation between the comparisons. Counting the control group twice is not acceptable.

Should I use odds ratios or risk ratios?

Risk ratios are easier to interpret, and odds ratios are convenient statistically and common for case-control data. The odds ratio overstates the risk ratio when the outcome is common. State the choice and the reason.

How should I handle studies with zero events?

Use methods that do not rely on adding a continuity correction, such as Mantel-Haenszel or a model fitted directly to the counts, and check the result with an alternative approach.

How many studies are needed?

There is no fixed minimum. Two studies can be combined mathematically, but with few studies the estimate of heterogeneity is imprecise and several checks lose power.

References

  1. Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  2. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
  3. Mantel N, Haenszel W. Statistical aspects of the analysis of data from retrospective studies of disease. J Natl Cancer Inst. 1959;22(4):719-748.
  4. Robins J, Breslow N, Greenland S. Estimators of the Mantel-Haenszel variance consistent in both sparse data and large-strata limiting models. Biometrics. 1986;42(2):311-323.
  5. Yusuf S, Peto R, Lewis J, Collins R, Sleight P. Beta blockade during and after myocardial infarction: an overview of the randomized trials. Prog Cardiovasc Dis. 1985;27(5):335-371.
  6. Bradburn MJ, Deeks JJ, Berlin JA, Russell Localio A. Much ado about nothing: a comparison of the performance of meta-analytical methods with rare events. Stat Med. 2007;26(1):53-77.
  7. Sweeting MJ, Sutton AJ, Lambert PC. What to add to nothing? Use and avoidance of continuity corrections in meta-analysis of sparse data. Stat Med. 2004;23(9):1351-1375.
  8. Stijnen T, Hamza TH, Ozdemir P. Random effects meta-analysis of event outcome in the framework of the generalized linear mixed model with applications in sparse data. Stat Med. 2010;29(29):3046-3067.
  9. IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. 2014;14:25.
  10. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.