The model
In the random-effects model each study i has its own true effect, written theta_i, which differs from study to study. The observed estimate y_i scatters around theta_i because of sampling error, with a variance v_i that is treated as known. The true effects themselves are assumed to scatter around an overall mean mu with a variance tau squared. Together this gives a two-level structure: sampling error within studies, and genuine differences between them. The average effect mu is the quantity of main interest, and tau squared measures how much the true effects vary.
The model has a clear interpretation. It does not assume that the studies are replicates of one experiment. It assumes that they are a sample from a population of possible studies that differ in participants, settings, delivery and design, and it asks what the average effect in that population is and how far the effect varies. This is the more realistic stance for most applied reviews, which is why random effects is so widely used. The page on the fixed-effect model describes the alternative, and the guide to the two models explains how to choose.
Weights and the pooled estimate
Under the random-effects model, the weight given to each study is the reciprocal of the sum of its within-study variance and the between-study variance, that is, 1 divided by the sum of v_i and tau squared. When tau squared is zero this reduces to the fixed-effect weight. When tau squared is large relative to the variances of the studies, the weights become nearly equal, and a large study loses much of its advantage over a small one. This is an inherent feature: if true effects genuinely differ, a very large study tells us precisely about its own effect, not necessarily about the average, so it should not dominate.
The pooled estimate is the weighted average using these weights, and its variance is 1 divided by their sum. Because the weights are smaller than under fixed effects, this variance is larger, and the confidence interval is wider. The width of the interval reflects both the uncertainty in each study and the uncertainty about how much the effects vary. This is why a random-effects interval is wider than a fixed-effect one whenever tau squared is above zero, and it is the main practical difference between the two analyses.
Estimating tau squared
The between-study variance must be estimated from the data, and several estimators exist. They give similar results with many large studies and can differ noticeably otherwise.
- DerSimonian-Laird
- A moment-based estimator that is simple, non-iterative and the default in much software. It is known to underestimate tau squared in many situations, especially with few studies or large heterogeneity, which makes intervals too narrow.
- Restricted maximum likelihood (REML)
- An iterative likelihood-based estimator with less bias than the moment estimator, widely recommended as a default for continuous outcomes.
- Paule-Mandel
- An iterative estimator that chooses tau squared so that the generalized Q statistic equals its expected value. It performs well, and is often recommended for binary outcomes.
- Maximum likelihood
- Efficient but biased downward with few studies, so less often preferred than REML.
- Bayesian estimators
- Produce a posterior distribution for tau squared, with the prior mattering when studies are few. See Bayesian meta-analysis.
Simulation studies have compared these estimators, and the choice should be stated and, where it could alter the conclusion, tested in a sensitivity analysis. Tau squared with few studies is very imprecise whichever estimator is used, and its confidence interval, which can be computed by profile likelihood or other methods, is often very wide. That width is an honest reflection of the information available.
Confidence intervals for the mean
The conventional random-effects interval treats the estimated tau squared as if it were known and uses the normal distribution. When studies are few, this ignores the uncertainty in tau squared and gives intervals that are too narrow, so that the nominal 95 percent interval covers the true mean less often than 95 percent of the time. The Hartung-Knapp-Sidik-Jonkman method addresses this by using a different variance estimate and a t distribution with k minus 1 degrees of freedom, where k is the number of studies. In simulations and in comparisons of published meta-analyses, it gives better coverage than the standard approach when studies are few and heterogeneity is moderate or large.
The adjustment has its own caveat: in some configurations, particularly with very homogeneous studies, the adjusted variance can be smaller than the conventional one, and a common remedy is to use whichever is larger. Many analysts now report the adjusted interval as a default for small numbers of studies and show the conventional one for comparison. The important point is not to rely on the conventional interval without thought when the number of studies is small.
The prediction interval
A confidence interval for the mean tells us where the average effect probably lies. It does not tell us what to expect in a particular new setting, because the effect there will differ from the average by an amount that depends on tau squared. The prediction interval addresses that question. It is the interval in which the true effect of a new study would be expected to fall, and it combines the uncertainty about the mean with the between-study variation: approximately the pooled estimate plus and minus a t value for k minus 2 degrees of freedom times the square root of tau squared plus the variance of the pooled estimate. It requires at least three studies, and it is wide and imprecise with few.
The prediction interval changes interpretation. A meta-analysis can show a clearly beneficial average effect, with a confidence interval well away from no effect, while the prediction interval includes no effect or even harm, which says that in some settings the intervention may not help. Reporting the prediction interval alongside the confidence interval is recommended by methodologists, and is increasingly expected. See prediction intervals.
Heterogeneity statistics
Several quantities describe heterogeneity, and they should be reported together and interpreted with care. Tau squared and its square root, tau, are on the scale of the effect measure and give the standard deviation of true effects, which is directly interpretable: if tau is 0.2 on the log risk ratio scale, the true effects typically vary by a factor of about 1.2 above and below the average. The I squared statistic expresses the proportion of observed variability that is due to heterogeneity and not chance, but it depends on the precision of the studies, and the same underlying variation yields a higher I squared in larger studies. Neither should be used with fixed thresholds to decide whether pooling is permitted. The guides on tau squared, I squared and heterogeneity explain them.
Few studies
Random-effects analysis with few studies, say fewer than five, is fragile. Tau squared is estimated with huge uncertainty, so the weights are unstable, the confidence interval is too narrow unless an adjustment is used, and the prediction interval is very wide or cannot be computed. The options are to use the adjusted interval, to present both fixed and random-effects results and explain why they differ, to use a Bayesian model with a weakly informative prior for heterogeneity, or to refrain from pooling and describe the studies individually. None of these removes the basic problem that the data contain little information about heterogeneity, and the report should say so. See Bayesian meta-analysis for the approach with few studies.
Practical guidance by number of studies
| Studies | Estimator and interval | Other points |
|---|---|---|
| Fewer than five | Adjusted interval; consider a Bayesian model with a weakly informative prior; show fixed-effect results too | Tau squared is very imprecise; a prediction interval is barely informative; consider not pooling |
| Five to ten | REML or Paule-Mandel with the Hartung-Knapp-Sidik-Jonkman interval | Report the prediction interval with its limits; avoid meta-regression with several moderators |
| More than ten | REML or Paule-Mandel; adjusted interval still reasonable | Tests for small-study effects become usable; subgroup analysis and single-moderator meta-regression are feasible |
These are starting points and not rules, and each analysis should be planned in the protocol and justified by the data and the question.
Reading tau on the original scale
The between-study standard deviation tau is useful when translated back to the scale readers understand. On the log scale of a ratio measure, a tau of 0.2 means that true effects typically vary by about 20 percent in either direction around the average, since the exponential of 0.2 is about 1.22. If the average risk ratio is 0.80, the true ratios in about two thirds of settings would then be expected to lie between roughly 0.65 and 0.98, on a normal approximation, which is a substantial range from clear benefit to barely any. A tau of 0.05 would give a range of about 0.76 to 0.84. On a standardized mean difference scale, a tau of 0.3 with an average of 0.4 implies that effects in some settings may be near 0.1 and in others near 0.7. Expressing heterogeneity in this way, instead of only as a percentage, makes its practical importance clear.
Criticisms and alternatives
The random-effects model is not free of problems. The assumption that the true effects are normally distributed may not hold, for example when there are two clusters of studies. The treatment of the within-study variances as known is an approximation, which is poor for small studies and sparse binary data. And a result that is dominated in weight by small studies can be affected by small-study effects, because random effects gives them more weight than the fixed-effect model does. Alternatives include models that use the exact likelihood for the data, such as binomial-normal models for binary outcomes, robust variance methods for dependent effects, and models for the distribution of effects that allow skewness or mixtures. Fixed-effect and random-effects results can be shown together as a check.
Reporting
A random-effects analysis is reported with the estimator of tau squared, the method for the confidence interval, the pooled estimate with its interval, the estimate of tau squared or tau, I squared with a note on its limits, and the prediction interval. The software, version and settings are given. The forest plot shows each study and the pooled estimate, with weights from the random-effects model. If the choice of estimator or interval method affects the conclusion, results under alternatives are shown in a sensitivity analysis, and the whole is reported according to PRISMA 2020.
How we can help
Support for random-effects meta-analysis
The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.
Feasibility check
A review of your studies and data to confirm that the method is suitable and which approach fits.
Analysis and figures
The analysis run to a prespecified plan, with forest plots and the other figures.
Methods and results text
Written for the manuscript and aligned with PRISMA 2020 or the relevant extension.
Manuscript and submission
Optional: the full paper, the reporting checklist and the submission materials.
Frequently asked questions
Why is a random-effects confidence interval wider than a fixed-effect one?
Because it includes the uncertainty about how much true effects vary between studies. When tau squared is above zero, the weights are smaller and more even, and the variance of the pooled estimate is larger.
Which estimator of tau squared should I use?
Restricted maximum likelihood and Paule-Mandel are generally preferred to DerSimonian-Laird in methodological comparisons. State the choice, and check it in a sensitivity analysis if it could change the conclusion.
What does the Hartung-Knapp adjustment do?
It uses a different variance estimate and a t distribution to give a confidence interval with better coverage when studies are few. Some analysts take the larger of the adjusted and conventional intervals.
Why report a prediction interval?
Because the confidence interval describes only the average effect. The prediction interval shows where the effect in a new setting would be expected to fall, and it can include no effect even when the average effect is clearly beneficial.
Can I use random effects with three studies?
It can be computed, but tau squared is very imprecise, intervals tend to be too narrow without adjustment, and a prediction interval is barely informative. Present the limits plainly and consider alternatives.
Is random effects always better than fixed effect?
Not always, but it is more realistic when studies differ, which is usual. A fixed-effect analysis can be justified when studies are very similar, and showing both is a useful check.
References
- DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials. 1986;7(3):177-188.
- Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.
- Veroniki AA, Jackson D, Viechtbauer W, et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Res Synth Methods. 2016;7(1):55-79.
- Langan D, Higgins JPT, Jackson D, et al. A comparison of heterogeneity variance estimators in simulated random-effects meta-analyses. Res Synth Methods. 2019;10(1):83-98.
- Paule RC, Mandel J. Consensus values and weighting factors. J Res Natl Bur Stand. 1982;87(5):377-385.
- Viechtbauer W. Bias and efficiency of meta-analytic variance estimators in the random-effects model. J Educ Behav Stat. 2005;30(3):261-293.
- Sidik K, Jonkman JN. A simple confidence interval for meta-analysis. Stat Med. 2002;21(21):3153-3159.
- IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. 2014;14:25.
- IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6(7):e010247.
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
- Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1-48.