Why rankings are attractive and risky
A network meta-analysis with eight or ten treatments produces dozens of pairwise comparisons, and few readers can absorb a league table of that size. A ranking looks like an answer: treatment D is first, A is last. It is easy to read, easy to quote and easy to misuse. A ranking compresses a set of uncertain estimates into an order, and the order can look firm when the estimates behind it are close together and imprecise.
Ranking metrics were introduced to address this need and are now reported in most network meta-analyses. They are useful when interpreted properly, and misleading when they are the only thing a reader looks at. The sections below show what they measure and what they leave out.
Rank probabilities
In a Bayesian network meta-analysis, each iteration of the model gives a set of treatment effects, and for that iteration the treatments can be ordered from best to worst. Counting how often each treatment takes each rank across many iterations gives rank probabilities: the probability that treatment A is first, second, third and so on. Frequentist analyses obtain similar quantities by resampling from the estimated distribution of effects.
The result is a rank matrix, with one row per treatment and one column per rank. Each row sums to one, because a treatment has to take some rank, and each column sums to one, because some treatment takes each rank. A bar chart called a rankogram displays one row. The probability of being best is the first column, and it is often what people quote, but it can be unstable, particularly for treatments studied in few small trials, because an imprecise treatment has a good chance of being best and also of being worst.
Calculating SUCRA
The cumulative ranking curve for a treatment plots the probability of being among the best b treatments, for b from 1 to the number of treatments minus one. SUCRA is the area under that curve, as a proportion of the maximum possible. With a treatments, SUCRA = (sum of the first a - 1 cumulative probabilities) / (a - 1). A treatment that is certainly best has SUCRA of 100 percent, and one that is certainly worst has 0 percent. A value of 50 percent means that the treatment is, on average, in the middle of the ranking.
A simulated four-treatment network illustrates the arithmetic. The rank probabilities are shown with the resulting SUCRA values and mean ranks.
| Treatment | Rank 1 | Rank 2 | Rank 3 | Rank 4 | SUCRA | Mean rank |
|---|---|---|---|---|---|---|
| A | 0.05 | 0.15 | 0.30 | 0.50 | 25% | 3.25 |
| B | 0.10 | 0.25 | 0.40 | 0.25 | 40% | 2.80 |
| C | 0.25 | 0.40 | 0.25 | 0.10 | 60% | 2.20 |
| D | 0.60 | 0.20 | 0.05 | 0.15 | 75% | 1.75 |
For treatment D the cumulative probabilities are 0.60, 0.80 and 0.85, which sum to 2.25, and 2.25 / 3 = 0.75, or 75 percent. Treatment D has the highest SUCRA, 75 percent, and a 60 percent chance of being first. But it also has a 15 percent chance of being last, so the distribution is bimodal. That pattern, a good chance of being best together with a real chance of being worst, is typical of a treatment with imprecise evidence, and it is hidden if only the SUCRA is shown. Treatment C, with a SUCRA of 60 percent, has a narrower spread. The SUCRA of A is 25 percent. The mean rank is a linear transformation of SUCRA, so the two carry the same information. The values are simulated.
P-scores and other metrics
The P-score is a frequentist counterpart of SUCRA, calculated from the point estimates and standard errors of pairwise comparisons, using the one-sided p values that treatment i is better than each other treatment. Rücker and Schwarzer showed that P-scores are numerically close to SUCRA values obtained from a Bayesian analysis of the same data under common assumptions, so the interpretation is the same: the average share of competitors that the treatment is better than, allowing for uncertainty.
Other ranking quantities have been suggested, including the probability of being best, the probability of being among the top two or three, the mean or median rank with its interval, and the probability that a treatment exceeds a clinically important threshold. Displays such as the rank heat plot and the Litmus rank-o-gram place several outcomes on a single chart, and the beading plot shows SUCRA values for multiple outcomes. Each makes a different trade-off between brevity and information.
What SUCRA does not tell you
SUCRA has real limitations, and several have been widely discussed.
- It ignores the size of the differences. A treatment can have a much higher SUCRA than another while the difference in effect is trivial or not clinically important. Ranks can be separated by tiny effects.
- It ignores certainty. Rankings rest on estimates whose certainty of evidence may be low or very low, particularly for comparisons that rely on indirect evidence. A high SUCRA from poor evidence is not a recommendation.
- It depends on the network. Adding or dropping a treatment changes the ranks of all others. Treatments with few trials tend to land at the extremes.
- It reduces a distribution to a number. Different rank probabilities can give identical SUCRA values, as the bimodal pattern above shows.
- It treats one outcome at a time. A treatment that is best for efficacy may be worst for safety, and a single ranking cannot capture the trade-off.
- It cannot reveal bias. If the network violates transitivity, or has inconsistency, the ranks inherit the problem.
Methodological papers have cautioned that rankings are often over-interpreted, and several have recommended that they be accompanied by the estimates and their certainty. Reviews of published network meta-analyses have found that some drew strong conclusions from SUCRA values without supporting estimates.
Rankings and certainty of evidence
Approaches that combine rankings with certainty have been developed. CINeMA and the GRADE approach to network meta-analysis rate the certainty of each estimate in the network, and one of their roles is to avoid statements such as "the best treatment" when the evidence is of low certainty. Some authors propose rankings based on the treatments for which the evidence has at least moderate certainty, and others suggest presenting the minimal important difference alongside estimates, so that readers can see which differences matter. The GRADE working group suggests presenting certainty-rated effect estimates, and using the ranks as supplementary information that is tied to those estimates. Whatever route is followed, a ranking that omits certainty is incomplete.
How to report rankings responsibly
A transparent report follows a few principles.
- Present effect estimates with confidence or credible intervals for the key comparisons first, usually in a league table or forest plot against a common comparator.
- Give the certainty of evidence for those estimates.
- Show rank probabilities or a rankogram, not only SUCRA, so that the uncertainty is visible.
- Report mean or median ranks with intervals when helpful.
- Say which model and software produced the ranks, with the number of iterations for Bayesian analyses.
- Rank each important outcome separately, and discuss trade-offs.
- Avoid wording such as "best treatment" unless the data justify it. Prefer wording like "treatment D had the highest probability of being among the most effective, but the evidence was of low certainty."
- Check that the network satisfies the assumptions of transitivity and consistency.
The PRISMA extension for network meta-analysis asks authors to describe how treatments were ranked and to give the results with uncertainty.
Linking ranks to clinically important differences
One way to make rankings more meaningful is to ask not where a treatment ranks, but how likely it is to beat the comparator by an amount that matters. If a minimal important difference is known for the outcome, the model can report the probability that each treatment exceeds that threshold relative to a reference. This answers a different question from SUCRA. Two treatments with the same SUCRA can differ greatly in the probability of a worthwhile benefit, and a treatment with a high rank can have a low probability of any important gain over standard care. Reporting both makes it easier to see when a ranking reflects real differences and when it reflects small differences in a crowded network. A practical check is to look at the league table beside the ranking: if the intervals for the first three treatments overlap heavily and all include the minimal important difference as a plausible value, then the order among them is not something a reader should act on, whatever the SUCRA values say. In such cases a more honest summary groups the top treatments together and says that the data cannot separate them, which is a far more useful statement for a guideline panel than a rank order that may not hold up once the next large trial is published.
Common mistakes
- Reporting only the treatment with the highest SUCRA as "the best".
- Showing SUCRA without the effect estimates or the certainty of evidence.
- Comparing SUCRA values between networks, which have different treatments and so different meaning.
- Ranking treatments for which only a few small trials exist alongside well-studied ones without comment.
- Ranking a single outcome and implying an overall verdict.
- Ignoring evidence of inconsistency or intransitivity when interpreting ranks.
- Presenting a Bayesian rank probability as if it were a statement of fact about the truth, and not conditional on the model and priors.
How we can help
We can fit frequentist and Bayesian network models, calculate rank probabilities, SUCRA and P-scores, create league tables, rankograms and certainty ratings, and write results that do not overstate the ranking. [OWNER VERIFICATION REQUIRED] The relevant services are network meta-analysis and statistical analysis.
Frequently asked questions
What does a SUCRA of 80 percent mean?
On average the treatment is better than about 80 percent of the other treatments, allowing for uncertainty. It does not say how large the difference is or how certain the evidence is.
Is the treatment with the highest SUCRA the best?
Not necessarily. Ranks can differ by trivial effects, depend on the network and rest on low-certainty evidence.
What is a P-score?
The frequentist analogue of SUCRA, calculated from point estimates and standard errors. Values are very close to SUCRA from a Bayesian analysis.
Why show rank probabilities as well as SUCRA?
Because different rank distributions can give the same SUCRA. A treatment with a good chance of being best and a real chance of being worst is hidden by the summary.
Should I rank treatments for safety and efficacy together?
Rank each outcome separately and discuss the trade-offs. A single ranking cannot capture them.
How should rankings relate to certainty of evidence?
Present certainty-rated effect estimates first. Rankings are supplementary and should not be used to recommend a treatment when the evidence is of low certainty.
References
- Salanti G, Ades AE, Ioannidis JPA. Graphical methods and numerical summaries for presenting results from multiple-treatment meta-analysis: an overview and tutorial. J Clin Epidemiol. 2011;64(2):163-171.
- Rucker G, Schwarzer G. Ranking treatments in frequentist network meta-analysis works without resampling methods. BMC Med Res Methodol. 2015;15:58.
- Mbuagbaw L, Rochwerg B, Jaeschke R, et al. Approaches to interpreting and choosing the best treatments in network meta-analyses. Syst Rev. 2017;6:79.
- Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082.
- Brignardello-Petersen R, Bonner A, Alexander PE, et al. Advances in the GRADE approach to rate the certainty in estimates from a network meta-analysis. J Clin Epidemiol. 2018;93:36-44.
- Hutton B, Salanti G, Caldwell DM, et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions: checklist and explanations. Ann Intern Med. 2015;162(11):777-784.
- Kibret T, Richer D, Beyene J. Bias in identification of the best treatment in a Bayesian network meta-analysis for binary outcome: a simulation study. Clin Epidemiol. 2014;6:451-460.