Indirect comparison: the basic idea
Suppose treatment B has been compared with A in some trials, and C has been compared with A in others, but no trial has compared B with C. An indirect comparison estimates B against C through the common comparator A. On the log scale, the indirect estimate of C versus B is the difference between the C-versus-A and B-versus-A estimates, and its variance is the sum of the two variances.
In a simulated example, the log risk ratio for B versus A is -0.30 (variance 0.02) and for C versus A is -0.10 (variance 0.03). The indirect estimate of C versus B is -0.10 - (-0.30) = 0.20, with variance 0.05 and standard error 0.224. As a risk ratio, that is 1.22 (95 percent CI 0.79 to 1.89). Notice that the interval is wider than either input, because uncertainty accumulates through the comparator.
This is the building block of every network meta-analysis. A network combines direct and indirect evidence across many treatments at once, and uses the same logic. The validity of the result depends on the assumptions below.
The transitivity assumption
Transitivity says that it is reasonable to compare B and C through A: if one could have randomized participants to A, B and C in a single trial, the B-versus-A and C-versus-A effects would apply to the same population and conditions. This requires that the trials of A versus B and A versus C are similar in all characteristics that modify the relative treatment effect, called effect modifiers. If, for example, the B-versus-A trials enrolled people with mild disease and the C-versus-A trials enrolled people with severe disease, and the effect of treatment depends on severity, then the indirect comparison confounds treatment with severity.
The assumption cannot be proved. It is a judgment based on clinical and methodological knowledge, and it is made at the level of the network, not of each trial. A review team has to decide which characteristics could modify treatment effects, record them for every trial, and compare their distribution across the comparisons in the network.
It is easy to confuse transitivity with heterogeneity. Heterogeneity is variation in the treatment effect among trials that make the same comparison. Transitivity concerns systematic differences in effect modifiers between trials that make different comparisons. A network can have low heterogeneity and still violate transitivity, and the reverse.
Assessing transitivity
Begin with the question of which characteristics are plausible effect modifiers. Typical candidates are disease severity, age, prior treatment, concomitant therapy, the dose and duration of treatments, the definition and timing of the outcome, setting and the year of the trial. The protocol should list them, with reasons. Then extract them for every trial, and compare their distributions across comparison groups. The table shows a simulated example for mean baseline severity on a 0 to 100 scale.
| Comparison | Mean baseline severity |
|---|---|
| A vs B (3 trials) | 42 |
| A vs C (2 trials) | 55 |
| B vs C (1 trial) | 38 |
The A-versus-C trials enrolled people with higher baseline severity (55) than the A-versus-B trials (42) and the single B-versus-C trial (38). If severity modifies the effect, indirect comparisons through A would be biased, and the review should say so. Approaches include a subgroup analysis or network meta-regression on the modifier, restriction of the network to a more homogeneous set of trials, or presenting the results with lower certainty. Graphs that display the distribution of each characteristic by comparison, and tables of trial characteristics, help others to judge.
Studies of published networks have found that transitivity is often not evaluated or reported, and that many authors rely on statistical tests of inconsistency alone. The PRISMA extension for network meta-analysis asks authors to describe how it was assessed.
Consistency and inconsistency
If transitivity holds, then the direct evidence for a comparison and the indirect evidence for the same comparison should agree, apart from chance. This is consistency, and the lack of it is inconsistency. It can only be assessed in loops of the network where both direct and indirect evidence exist.
Continue the example. Suppose a trial did compare C with B directly, with a log risk ratio of 0.05 (variance 0.04), which is a risk ratio of 1.05 (95 percent CI 0.71 to 1.56). The indirect estimate was 0.20. The difference is -0.15, with a standard error of sqrt(0.05 + 0.04) = 0.30, so the test statistic is z = -0.50. That is far from the 1.96 threshold, so there is no statistical evidence of inconsistency in this loop. Combining the two with inverse-variance weights gives a log risk ratio of 0.12.
The lack of significance should not be read as proof of consistency. The standard error of the difference is 0.30, so the smallest difference this loop could detect with 80 percent power is about 0.84 on the log scale, which is more than four times the size of the indirect estimate itself. Tests for inconsistency are known to have low power, particularly in networks with few trials per comparison.
Statistical methods for inconsistency
Several approaches are available, and they answer slightly different questions.
- Loop-specific approach. Compares direct and indirect estimates in each closed loop, as in the example. Simple and transparent, but limited to the loops.
- Node-splitting. For each comparison with both sources of evidence, separates direct and indirect estimates and tests the difference, using the whole network to produce the indirect estimate. It is available in Bayesian and frequentist software.
- Design-by-treatment interaction model. A global test of inconsistency across the whole network, which also handles multi-arm trials correctly. It gives a single test, with low power, and no information on where the inconsistency lies.
- Side-splitting and related approaches. Variations on the same idea.
- Comparison of consistency and inconsistency models. A model fit approach, using deviance or information criteria, for Bayesian analyses.
Because the tests have low power, a non-significant result is weak reassurance. A significant result needs investigation: check the data for extraction errors, look at the effect modifiers in the loop, and consider whether a particular trial is an outlier. Inconsistency and heterogeneity are linked, since high heterogeneity within comparisons makes inconsistency harder to detect, and the models should be fitted with attention to both.
What to do if the assumptions look doubtful
When transitivity looks doubtful, or inconsistency is found, options include the following.
- Investigate the cause. Review the data and the trial characteristics, and look for a modifier that explains the discrepancy.
- Adjust. Network meta-regression, with caution about the number of trials per covariate, can model an effect modifier. Subgroup or sensitivity analyses can restrict the network.
- Split the network. Analyse separate, more homogeneous networks, or exclude a comparison or a trial that causes the problem, with a stated rationale.
- Use only direct evidence. For the comparisons affected, rely on direct estimates and describe the indirect ones as unreliable.
- Lower the certainty. In GRADE and CINeMA, intransitivity and incoherence are reasons to lower confidence in the network estimates.
- Be open about it. Report the assessment, the findings and the effect on conclusions.
It is better to define the network carefully at the protocol stage than to fix problems afterwards. The question of which treatments are eligible to join the network, and which should be left out as not jointly randomizable, is the first line of defense. If two treatments could never realistically be offered to the same patient population, comparing them through a hub treatment makes little sense.
Reporting
The report should state which effect modifiers were considered and why, show their distribution by comparison, describe the evaluation of transitivity, give the results of the local and global inconsistency assessments, and say how they affected the analysis and the certainty. Display the network graph with the numbers of trials and participants for each comparison, since a thin comparison deserves attention. Provide a table giving, for every comparison with both sources, the direct, indirect and network estimates. PRISMA-NMA asks for the methods to assess inconsistency and the results.
Thinking about which networks are credible
Some networks are inherently more credible than others. A network of treatments within one drug class, tested in similar patients against a common placebo, is likely to be nearly transitive. A network that links a surgical procedure, a drug and a behavioral program through a shared "usual care" arm is much less likely to be, since "usual care" means different things in each setting and the populations differ. The geometry matters too. A star-shaped network, in which every treatment was compared only with one common comparator, has no loops, so consistency cannot be examined at all, and every comparison between active treatments rests entirely on the assumption of transitivity. A network with many closed loops offers more checks, though the checks have limited power.
Multi-arm trials need special care. A three-arm trial contributes correlated comparisons, and treating them as independent two-arm trials inflates their weight. Network software handles this correctly, and the analyst should confirm that it does. It also matters that a multi-arm trial is internally consistent by design, so it adds no inconsistency to the network, although it can mask inconsistency elsewhere.
Common pitfalls
- Testing for inconsistency and skipping the assessment of transitivity.
- Reading a non-significant test as proof of consistency.
- Choosing effect modifiers after seeing the results.
- Including treatments that are not realistic alternatives for the same population, only because a trial exists.
- Ignoring differences in dose, timing or outcome definition across comparisons.
- Treating a star-shaped network as if it had been checked for consistency.
- Using the fixed-effect model without considering the heterogeneity that blurs inconsistency.
How we can help
We can define the network, choose and extract effect modifiers, judge transitivity, run node-splitting and global tests, fit network meta-regression, rate certainty with CINeMA or GRADE and write the assessment so that readers can follow it. [OWNER VERIFICATION REQUIRED] The relevant services are network meta-analysis and statistical analysis.
Frequently asked questions
What is transitivity?
The assumption that trials comparing different treatments are similar enough in effect modifiers for an indirect comparison to be valid. It is a judgment, not a test.
What is inconsistency?
A disagreement between direct and indirect estimates of the same comparison, greater than chance explains.
Can I test transitivity statistically?
Not directly. It is assessed by comparing the distribution of effect modifiers across comparisons, using clinical knowledge.
Does a non-significant inconsistency test mean the network is consistent?
No. These tests have low power, especially with few trials, so a non-significant result is weak reassurance.
What is node-splitting?
A method that separates direct from indirect evidence for each comparison, and tests the difference, using the rest of the network for the indirect estimate.
What if I find inconsistency?
Check the data, look for effect modifiers, consider network meta-regression or splitting the network, and lower the certainty of affected estimates.
References
- Salanti G. Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, many concerns for the next generation evidence synthesis tool. Res Synth Methods. 2012;3(2):80-97.
- Dias S, Welton NJ, Caldwell DM, Ades AE. Checking consistency in mixed treatment comparison meta-analysis. Stat Med. 2010;29(7-8):932-944.
- Higgins JPT, Jackson D, Barrett JK, Lu G, Ades AE, White IR. Consistency and inconsistency in network meta-analysis: concepts and models for multi-arm studies. Res Synth Methods. 2012;3(2):98-110.
- Bucher HC, Guyatt GH, Griffith LE, Walter SD. The results of direct and indirect treatment comparisons in meta-analysis of randomized controlled trials. J Clin Epidemiol. 1997;50(6):683-691.
- Jansen JP, Naci H. Is network meta-analysis as valid as standard pairwise meta-analysis? It all depends on the distribution of effect modifiers. BMC Med. 2013;11:159.
- Cipriani A, Higgins JPT, Geddes JR, Salanti G. Conceptual and technical challenges in network meta-analysis. Ann Intern Med. 2013;159(2):130-137.
- Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082.
- Hutton B, Salanti G, Caldwell DM, et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions: checklist and explanations. Ann Intern Med. 2015;162(11):777-784.