Two different questions
A random-effects meta-analysis reports an average effect and an interval around it. The interval is usually a confidence interval, and it answers one question: how precisely do these studies tell us the average effect? It narrows as studies are added. But when the true effect differs from study to study, a second question matters to anyone who will apply the result: what effect should I expect in my own setting, which is not one of the studies? The average does not answer that, because the effect in any particular setting will differ from the average by an amount that depends on how much true effects vary.
The prediction interval addresses the second question. It combines the uncertainty about the average with the variation between studies, and gives a range in which the effect in a new study would be expected to lie, with a stated probability. Because it includes the extra source of variation, it is wider than the confidence interval, sometimes much wider. It tells the reader whether the effect is likely to be consistent across settings, or whether it might sometimes be absent or reversed. The page on random-effects meta-analysis describes the model behind it.
How it is calculated
The approximate 95 percent prediction interval is the pooled estimate plus and minus a critical value times the square root of the sum of two variances: the between-study variance, tau squared, and the variance of the pooled estimate. The critical value comes from a t distribution with k minus 2 degrees of freedom, where k is the number of studies, which is why at least three studies are needed, and why the interval is wide with few. The larger the between-study variance, the wider the interval, and with no heterogeneity it reduces to nearly the confidence interval, widened only by the use of the t distribution.
The formula is an approximation that relies on tau squared being estimated well and on the effects being normally distributed. With few studies the estimate of tau squared is unreliable, and so is the interval. Bayesian methods give a predictive distribution that accounts for the uncertainty in tau squared directly, and is often preferred when studies are few. Software commonly offers the interval as an option, and its availability has made routine reporting easy. See Bayesian meta-analysis.
A worked example
Eight simulated studies report standardized mean differences. They are invented for illustration and are not real research. A random-effects analysis gives a pooled estimate of 0.38, with a standard error of 0.071, so the 95 percent confidence interval runs from 0.25 to 0.52. The estimated between-study variance is 0.0204, so tau is 0.14. The prediction interval, using a t distribution with 6 degrees of freedom, runs from -0.01 to 0.77.
| Study | Effect (SMD) | Standard error |
|---|---|---|
| Study A | 0.55 | 0.14 |
| Study B | 0.20 | 0.10 |
| Study C | 0.70 | 0.18 |
| Study D | 0.35 | 0.12 |
| Study E | 0.05 | 0.16 |
| Study F | 0.48 | 0.11 |
| Study G | 0.62 | 0.20 |
| Study H | 0.28 | 0.13 |
| Pooled (random effects) | 0.38 | 0.071 |
The two intervals tell different stories. The confidence interval, from 0.25 to 0.52, lies entirely above zero, so the average effect is clearly positive. The prediction interval, from -0.01 to 0.77, includes zero and extends below it, so in a new setting the true effect could be anywhere from slightly negative to large and positive. A reader who looked only at the confidence interval would conclude that the intervention works, with some confidence. A reader who also saw the prediction interval would conclude that it works on average, and that the effect is variable enough that it might not in a particular place. Both are true, and the second is more useful to someone deciding whether to adopt the intervention locally.
Interpreting it
The prediction interval is not a statement about the participants in a study, and it is not a range for individual patients. It concerns the true effect, the average effect in a population of participants, in a new study. It does not shrink as studies accumulate, because it contains the true variation between settings, which more studies estimate more precisely but do not remove. Several readings are common. If the whole interval lies on the beneficial side, the effect is likely to be beneficial in any setting similar to those studied. If it straddles no effect, the effect varies enough that benefit cannot be assured everywhere. If it is very wide, the evidence is too uncertain to say much about a new setting. Each reading should be tied to the question: for a decision in a particular setting, one wants to know whether the setting resembles those in which the effect was larger or smaller, which is why exploration of heterogeneity matters.
The prediction interval is also a useful check on the claim that an intervention has a consistent effect. A review that describes an effect as consistent, with a prediction interval that includes no effect, has not shown consistency.
Relation to I-squared and tau-squared
The prediction interval complements the other heterogeneity measures. I squared gives the share of variability due to heterogeneity, as a percentage, without telling the reader what the heterogeneity means for the effect. Tau squared, and its square root tau, give the variation on the scale of the effect, which can be hard to translate. The prediction interval does the translation, because it converts heterogeneity and uncertainty into a range of effects. In practice a reader can interpret a prediction interval directly, while the meaning of an I squared of 70 percent is far from obvious. For this reason methodologists have argued that a prediction interval should accompany every random-effects meta-analysis. See the guides on I squared and tau squared.
With few studies
The interval is hardest to use when it is most needed. With three or four studies, the t distribution has one or two degrees of freedom and its critical value is very large, so the interval is enormous, and the estimate of tau squared is so imprecise that the interval says little. This is not a defect of the method but a reflection of how little the data say about variation across settings. In such cases the options are to report the interval and describe it as uninformative, to use a Bayesian model with a justified prior for heterogeneity, to reason about heterogeneity from the clinical and methodological features of the studies, or to explain that the evidence is insufficient to say whether effects vary. Reporting that the interval is wide is a more honest summary than omitting it.
Using the interval in a decision
A decision maker deciding whether to adopt an intervention has to ask what is likely to happen in their own setting. The prediction interval supplies the range of effects among settings like those studied, but it cannot say where a given setting falls inside it. Three questions help. Does the interval lie wholly on the beneficial side of the threshold that matters? If so, adoption is supported across the range. Does it straddle the threshold? Then the effect depends on the setting, and the review of the characteristics associated with larger and smaller effects, from subgroup analysis or meta-regression, becomes essential. And is it so wide that it carries little information? Then the evidence does not yet support a confident statement for new settings, and a local evaluation may be worth doing. Used in this way, the interval turns heterogeneity from a technical statistic into a practical input.
Predictive distributions in Bayesian analysis
In a Bayesian meta-analysis the corresponding quantity is the posterior predictive distribution of the effect in a new study. It is obtained by drawing, for each simulated value of the overall mean and the between-study variance from their joint posterior, a new study effect from the distribution those values imply. Because it integrates over the uncertainty in both parameters, it avoids the approximation involved in the t-based interval, and it is straightforward to produce from the simulation output. From it one can read not only an interval but the probability that the effect in a new study exceeds a threshold, which answers directly the question of how likely the intervention is to be beneficial in a new setting. The result depends on the prior for the heterogeneity when studies are few, so a sensitivity analysis to that prior is reported. See Bayesian meta-analysis.
Showing it on a forest plot
Many software packages can add the prediction interval to a forest plot, as a line through or beneath the diamond, or a pale band around it. The pooled estimate is shown as the diamond, with the confidence interval as its width, and the prediction interval as a longer line extending beyond it. The picture makes the difference between the two immediately visible. The plot should label the line, since readers will not always recognize it, and it should state the number of studies it is based on. See how to read a forest plot.
Reporting
A report should state the method used to compute the prediction interval, the estimator of tau squared and the degrees of freedom, give the interval in the text and tables alongside the confidence interval, and interpret it. It should say how many studies it is based on. In the discussion, the interpretation should draw out its consequence for applying the result: consistent benefit, benefit on average with variation, or insufficient information. In GRADE, a prediction interval that crosses the threshold of importance, while the confidence interval does not, is relevant to the judgment of inconsistency and imprecision. See GRADE.
Common mistakes
- Reporting only the confidence interval of a random-effects analysis and describing the effect as consistent.
- Confusing the prediction interval with a range for individual participants.
- Computing it with two studies, or interpreting a very wide interval from three as informative.
- Using it under a fixed-effect model, where it is not defined in the same way.
- Expecting it to narrow as more studies are added.
- Omitting the method, so that readers cannot tell how it was computed.
Support
Prediction intervals are reported as part of the meta-analysis service, and Bayesian predictive distributions as part of the Bayesian meta-analysis service.
Get a quoteSend your data and target journal.
Frequently asked questions
What is the difference between a confidence interval and a prediction interval?
A confidence interval shows the uncertainty about the average effect. A prediction interval shows the range in which the true effect of a new study would be expected to fall, including the real variation between studies, so it is wider.
How many studies are needed to compute a prediction interval?
At least three, because the t distribution needs at least one degree of freedom, and many more for the interval to be informative.
Can the prediction interval include zero when the confidence interval does not?
Yes. That is the case in which the average effect is clearly beneficial but the effect in some settings could be absent or reversed.
Does a prediction interval apply to individual patients?
No. It concerns the true average effect in a new study, not the response of an individual.
Why does the prediction interval not shrink with more studies?
Because it includes real variation between studies. More studies estimate that variation more precisely but do not remove it.
Should I always report it?
In random-effects meta-analyses with three or more studies, yes, with an interpretation, since it conveys what the heterogeneity means for the effect.
References
- IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6(7):e010247.
- Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.
- Riley RD, Higgins JPT, Deeks JJ. Interpretation of random effects meta-analyses. BMJ. 2011;342:d549.
- Partlett C, Riley RD. Random effects meta-analysis: coverage performance of 95% confidence and prediction intervals following REML estimation. Stat Med. 2017;36(2):301-317.
- Nagashima K, Noma H, Furukawa TA. Prediction intervals for random-effects meta-analysis: a confidence distribution approach. Stat Methods Med Res. 2019;28(6):1689-1702.
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
- Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.