Method

Network meta-analysis: the method

Network meta-analysis compares several interventions in a single model, combining direct evidence from head-to-head trials with indirect evidence through common comparators. This page describes the structure of the method, how its models work, the assumptions on which it rests and how its results are summarized.

The central idea

Suppose trials have compared treatment A with placebo and treatment B with placebo, but none has compared A with B. A pairwise analysis can say nothing about A against B. A network meta-analysis can: if A reduces the outcome by a certain amount relative to placebo, and B by another, the difference between those amounts is an indirect estimate of A against B. The method generalizes this idea to a whole network of treatments, estimating all relative effects consistently and using every trial. When both direct and indirect evidence are available for a comparison, they are combined into a mixed estimate that is more precise than either.

The indirect estimate of A against B through a common comparator C is the difference between the A-versus-C and B-versus-C estimates, and its variance is the sum of their variances. It is therefore always less precise than direct evidence of the same quantity from a trial of similar size. The value of the method lies in its ability to use all the evidence at once, to say something about comparisons that have never been tested, and to rank the options. Its risk lies in the assumptions needed for the indirect evidence to be valid. A practical outline for clients is in the network meta-analysis service.

Network structure

A network is described by a graph in which nodes are treatments and edges are direct comparisons. Several features matter. Connectedness is essential: every treatment must be linked to every other by some path, otherwise the network splits into separate sub-networks that cannot be compared. The geometry, that is, which comparisons exist and how many trials inform each, determines where information comes from. A star-shaped network, in which all treatments have been compared with one reference such as placebo, supports only indirect comparisons among the active treatments. Loops, formed when three or more treatments are all compared with each other in different trials, allow the consistency of direct and indirect evidence to be examined. Designs are the sets of treatments that appear together in a trial, and the pattern of designs affects the analysis.

The network graph, with node size reflecting the number of participants and edge thickness the number of trials, is the standard first display. It shows at a glance where the evidence is dense and where it is sparse, and it should be shown in every report.

Transitivity and similarity

The key assumption is transitivity. For an indirect comparison of A and B through C to be valid, the trials of A against C and of B against C must be similar in every characteristic that modifies the relative effects, such as the severity of illness, the age of patients, the dose, the duration of follow-up and the definition of the outcome. The assumption can be stated as: the participants in the A-versus-C trials could in principle have been randomized to B. If the trials making different comparisons differ in effect modifiers, the indirect estimates are biased, and no statistical adjustment can fix this unless the modifiers are measured and modeled.

Transitivity cannot be tested directly. It is examined by comparing the distribution of potential effect modifiers across the comparisons, using tables and plots of study characteristics, and by judging whether any imbalance is large enough to matter. This work belongs at the protocol stage, where the effect modifiers are specified in advance, and at the data stage, where they are extracted. Differences that cannot be resolved can limit the network, or lead to the exclusion of some treatments or trials, or to an adjusted analysis with meta-regression. See transitivity and inconsistency.

Models

Two modeling approaches are used. The contrast-based model, the most common, takes as its data the relative effect of each treatment against a reference treatment in each trial, with a covariance structure for multi-arm trials, and models the true relative effects as linear combinations of a set of basic parameters, the effects of each treatment against a chosen reference. All other contrasts follow from these parameters through the consistency equations, which state that the effect of B against A equals the effect of B against C minus the effect of A against C. In the random-effects version, the trial-specific relative effects vary around these means with a between-trial variance, commonly assumed to be the same for all comparisons.

The arm-based model works with the outcome in each arm, for instance the number of events and participants, and models the treatment-specific outcomes in each trial directly, typically with a trial-specific baseline. It can use all arm-level information and handles sparse data with exact likelihoods, but its treatment of baseline risk as random across trials has been debated, because it can break the protection that randomization gives within each trial. Most guidance recommends the contrast-based formulation as the default, with arm-based approaches reserved for particular purposes.

The models can be fitted in frequentist or Bayesian frameworks. Frequentist implementations include multivariate meta-analysis models and graph-theoretical methods that are quick and give confidence intervals. Bayesian implementations are fitted by simulation and give direct probability statements. The Bayesian meta-analysis page describes that framework.

Multi-arm trials

A trial with three or more arms contributes several comparisons that share arms, so their estimates are correlated. If they are entered as independent two-arm comparisons, the shared arm is double-counted and the precision is overstated. Network meta-analysis handles this directly by modeling the multi-arm trial as a unit, with a covariance matrix between its contrasts. Contrast-based software handles this automatically given the right data format, which is why the data are usually prepared with one row per comparison and a record of the arm structure. Treating multi-arm trials correctly is one of the main reasons to use network meta-analysis and not a collection of pairwise analyses.

Consistency and inconsistency

Consistency is the statistical agreement of direct and indirect evidence. In a network with closed loops, it can be checked. A loop-based approach compares the direct and indirect estimates in each loop and computes the inconsistency factor. A node-splitting approach separates direct from indirect evidence for each comparison in turn and tests their difference. Global approaches fit a model that allows inconsistency, such as the design-by-treatment interaction model, and compare it with the consistency model, giving an overall test. These approaches have low power, and their failure to find inconsistency should not be read as showing that the network is consistent.

When inconsistency is found, the response is to explore its sources, which are usually differences in effect modifiers, and not simply to report a number. Options include examining and adjusting for the modifiers by meta-regression, excluding trials that are clearly different, splitting nodes that combine different interventions, or using a model that allows for inconsistency while acknowledging that the estimates then have a different meaning.

Ranking and presentation

The results are presented as a league table that shows the estimated effect, with its interval, for every pair of treatments, along with forest plots of each treatment against a reference. Rankings are summarized with the surface under the cumulative ranking curve and the P-score, which give for each treatment a number between zero and one. These summaries should be interpreted with caution. A treatment can be ranked first while the evidence that it is better than the second is weak, and rankings depend on the treatments included and ignore the certainty of the evidence. Many methodologists advise presenting rankings only with the relative effects, their intervals and the certainty ratings, and not as a stand-alone result. See SUCRA and rankings.

Certainty and reporting bias

The certainty of the evidence is assessed for each comparison, using an approach that extends GRADE to networks, because it varies with the contributions of direct and indirect evidence and with the risk of bias and the other usual domains for each. Small-study effects are examined with comparison-adjusted funnel plots, which orient all the comparisons in a common direction, and with network meta-regression on study variance. As in pairwise analysis, these assessments are less powerful when comparisons have few trials.

Reporting

A network meta-analysis is reported against PRISMA 2020 and its extension for network meta-analyses. The report describes the network and its geometry, the way in which treatments were grouped into nodes, the assessment of transitivity, the model and its assumptions, the software and code, the assessment of inconsistency and small-study effects, the results for all comparisons and the certainty of the evidence. The code and data are made available so that the analysis can be reproduced. See PRISMA extensions.

A small illustration

Consider three treatments, A, B and placebo P, with trials of A against P and of B against P only. If the pooled log risk ratio for A against P is minus 0.40 with variance 0.01, and for B against P is minus 0.25 with variance 0.02, the indirect estimate of A against B is the difference, minus 0.15, and its variance is the sum, 0.03, giving a standard error of about 0.17. On the risk ratio scale this is about 0.86, with a 95 percent interval of roughly 0.61 to 1.21. The estimate favors A, but the interval includes no difference, and it is wider than either of the two direct estimates, because the uncertainty from both has been added. If a direct trial of A against B existed, its estimate would be combined with this indirect one, weighted by precision, and the mixed estimate would be more precise than either. The numbers here are chosen to illustrate the arithmetic and are not real data.

Software and workflow

Network meta-analysis can be run in R with packages for frequentist models, such as netmeta, and for Bayesian models, using general-purpose engines such as Stan or JAGS or packages built on them, and in Stata with a set of network commands. Whatever the program, the workflow is similar: prepare the data in the format the software expects, with one row per comparison or per arm; draw and check the network graph; fit the consistency model; assess heterogeneity and inconsistency; produce the relative effects, rankings and plots; and run the sensitivity analyses. The data preparation, especially the coding of treatments and multi-arm trials, is where mistakes most often occur, so it is checked and documented, and the scripts are kept so that the analysis can be reproduced.

Limitations

The method depends on assumptions that cannot be verified from the data, notably transitivity. Indirect evidence is observational in nature, as the participants were not randomized between the treatments compared indirectly. Sparse networks produce imprecise estimates and weak tests of inconsistency. Rankings are easily over-interpreted. The grouping of treatments into nodes involves judgments that affect the results. And networks that look large can rest on few trials per comparison. For these reasons a network meta-analysis supports decisions and does not replace judgment, and its results are not advice about the care of an individual patient.

Network meta-analyses should be planned with the assumptions in mind from the start, with effect modifiers named in the protocol.

How we can help

Support

Support for network meta-analysis

The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.

Get a quoteTell us your question, the number of studies and your target journal.

Frequently asked questions

What is an indirect comparison?

An estimate of the difference between two treatments obtained through a common comparator. If A and B have each been compared with C, the difference between the A-versus-C and B-versus-C effects estimates A against B.

What is the consistency assumption?

That direct and indirect evidence for a comparison agree, apart from chance. It can be checked in loops of the network, with tests of low power, and depends on the transitivity assumption.

Do I need a loop in my network?

No. A network without loops can still be analyzed, but consistency cannot then be checked statistically, so the transitivity assumption has to be argued on clinical grounds.

Should I use contrast-based or arm-based models?

The contrast-based approach is the usual default because it preserves the within-trial randomization. Arm-based models are used for particular purposes, with their different assumptions stated.

Why does the grouping of treatments matter?

Because each node should represent a single, clearly defined intervention. Lumping different doses or combinations can hide differences and break transitivity, while excessive splitting leaves the network too sparse.

Can I rely on SUCRA values?

Use them only with the relative effects, their intervals and the certainty of the evidence. A high ranking can reflect weak evidence, and rankings change with the treatments included.

References

  1. Hutton B, Salanti G, Caldwell DM, et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions. Ann Intern Med. 2015;162(11):777-784.
  2. Bucher HC, Guyatt GH, Griffith LE, Walter SD. The results of direct and indirect treatment comparisons in meta-analysis of randomized controlled trials. J Clin Epidemiol. 1997;50(6):683-691.
  3. Lu G, Ades AE. Combination of direct and indirect evidence in mixed treatment comparisons. Stat Med. 2004;23(20):3105-3124.
  4. Salanti G. Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, many concerns for the next generation evidence synthesis tool. Res Synth Methods. 2012;3(2):80-97.
  5. Higgins JPT, Jackson D, Barrett JK, Lu G, Ades AE, White IR. Consistency and inconsistency in network meta-analysis: concepts and models for multi-arm studies. Res Synth Methods. 2012;3(2):98-110.
  6. White IR, Barrett JK, Jackson D, Higgins JPT. Consistency and inconsistency in network meta-analysis: model estimation using multivariate meta-regression. Res Synth Methods. 2012;3(2):111-125.
  7. Dias S, Welton NJ, Caldwell DM, Ades AE. Checking consistency in mixed treatment comparison meta-analysis. Stat Med. 2010;29(7-8):932-944.
  8. Dias S, Sutton AJ, Ades AE, Welton NJ. Evidence synthesis for decision making 2: a generalized linear modeling framework for pairwise and network meta-analysis of randomized controlled trials. Med Decis Making. 2013;33(5):607-617.
  9. Rucker G. Network meta-analysis, electrical networks and graph theory. Res Synth Methods. 2012;3(4):312-324.
  10. Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Med. 2013;11:159.
  11. Chaimani A, Salanti G. Using network meta-analysis to evaluate the existence of small-study effects in a network of interventions. Res Synth Methods. 2012;3(2):161-176.
  12. Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.