Guide

What is tau-squared (τ²)?

Tau-squared is the variance of the true effects across studies in a random-effects meta-analysis. It is the quantity that makes random effects different from fixed effect, and it has to be estimated. This guide explains what it means, compares the main estimators on one dataset, and shows why its uncertainty matters.

What tau-squared is

A random-effects meta-analysis assumes that each study has its own true effect, and that these true effects are scattered around an average. Tau squared is the variance of that scatter. If it is zero, all studies share one true effect and the random-effects analysis collapses to the fixed-effect analysis. The larger it is, the more the true effects differ, and the more the random-effects weights depart from the fixed-effect ones. It is measured in the squared units of the effect measure: on the log risk ratio scale it is a variance of log ratios, and for standardized mean differences it is in squared standard deviation units.

Because the square is hard to picture, its square root, tau, is easier to interpret. It is the standard deviation of the true effects, in the units of the effect measure. If tau is 0.2 on the log risk ratio scale, true effects typically differ from the average by a factor of about 1.2, up or down. That interpretation, rather than the raw variance, is what a reader needs. The model in which tau squared appears is described on the page for random-effects meta-analysis.

What it does in the analysis

Tau squared enters the analysis in three places. First, it adjusts the weights: each study's weight is the reciprocal of its own variance plus tau squared, so a larger tau squared makes the weights more equal, because the between-study variance adds the same amount to every study's total variance. Second, it widens the confidence interval, since the variance of the pooled estimate is the reciprocal of the sum of the weights, and smaller weights give a larger variance. Third, it is a component of the prediction interval, which is the pooled estimate plus and minus a critical value times the square root of tau squared plus the variance of the pooled estimate. The description of heterogeneity in a report often states tau squared, tau and I squared together. See prediction intervals.

Estimators

DerSimonian-Laird (DL)
A moment estimator: tau squared is the excess of Q over its degrees of freedom, divided by a scaling constant derived from the weights. It is explicit and fast, the default in much software and the standard in older literature. It is known to underestimate tau squared when studies are few or heterogeneity large.
Restricted maximum likelihood (REML)
An iterative likelihood-based estimator that corrects for the loss of a degree of freedom in estimating the mean. It is less biased than maximum likelihood and generally recommended as a default for continuous outcomes.
Paule-Mandel (PM)
An iterative estimator that chooses tau squared so that the generalized Q statistic equals its expected value, its degrees of freedom. It performs well, and is often recommended for binary data.
Maximum likelihood (ML)
An iterative estimator that is biased downward in small samples, so less often preferred than REML.
Sidik-Jonkman and others
Estimators with other properties, some positively biased, sometimes used as alternatives or in sensitivity analyses.

All of them truncate at zero, because a variance cannot be negative, so an estimate of zero means only that the data show no excess variation, not that the true effects are identical.

A worked comparison of three estimators

Eight simulated studies, invented for illustration, report standardized mean differences with standard errors from 0.10 to 0.20. The Q statistic is 15.39 on 7 degrees of freedom. The three estimators give different answers:

Estimates of tau-squared for the same data
EstimatorTau-squaredTau
DerSimonian-Laird0.02350.15
Paule-Mandel0.02730.17
REML0.02440.16

In this dataset the three estimates are fairly close: they range from 0.0235 to 0.0273, with Paule-Mandel at 1.16 times the DerSimonian-Laird value and REML at 1.04 times. That is typical when there are about eight studies and moderate heterogeneity. The estimators diverge more when studies are few, when heterogeneity is large, or when study sizes are very unequal, and simulation studies find that the DerSimonian-Laird estimate tends to be the lowest in those conditions. Even here, the choice changes the weights slightly and the prediction interval a little, so the estimator should be stated, and, if the conclusion could depend on it, the others shown as a sensitivity analysis.

When the estimators diverge

The comparison above is a mild case. In collections with three to five studies, or with a very large and a few very small studies, the differences can be much greater, and in some configurations one estimator returns zero while another returns a substantial value. This is because the estimators weigh the information differently: the moment estimator takes the excess of Q over its expectation under the fixed-effect weights, which are dominated by the largest study, while the iterative estimators update the weights as tau squared changes. Simulation studies comparing a dozen estimators have found that no one is best everywhere, that DerSimonian-Laird is often biased low, and that REML and Paule-Mandel are good general-purpose choices, with the latter often preferred for binary outcomes. These findings are the reason that methodological guidance now advises against using the DL estimator as an unexamined default, while accepting it as a common baseline for comparison with older reviews.

Uncertainty in tau-squared

Tau squared is estimated with considerable uncertainty. With a modest number of studies, its confidence interval typically spans from zero to several times the point estimate, which means that the amount of heterogeneity is poorly known. Several methods give such intervals. The Q-profile method inverts the generalized Q statistic and has good coverage in many situations. Profile likelihood intervals are available with REML. The Biggerstaff-Jackson method is another. Reporting the interval, or at least acknowledging the uncertainty, is good practice and is increasingly expected. It also helps to guard against the habit of treating the point estimate as if it were known, which is what the conventional confidence interval for the pooled effect does, with the result that the interval is too narrow when studies are few. The Hartung-Knapp adjustment addresses that. See random-effects meta-analysis.

Interpreting tau on the original scale

The practical meaning of tau depends on the effect measure. For a ratio measure analyzed on the log scale, the exponential of tau is the typical factor by which true effects differ from the average. A tau of 0.1 means about 10 percent differences, which would often be unimportant, whereas a tau of 0.5 means that effects typically differ by a factor of 1.65, which spans a range from clear benefit to little or even harm around a moderate average. For standardized mean differences, a tau of 0.3 means that true effects typically lie about a third of a standard deviation above or below the average, a difference between small and large effects by conventional labels. Judging whether a given tau is important requires a threshold of what difference would matter in practice, and that is a decision for the subject expert, not the statistician. Translating tau into such terms is much more useful than quoting tau squared.

With few studies

The estimate of tau squared is unreliable when the meta-analysis has fewer than about five studies, and in that situation several problems appear. The estimate may be truncated at zero by chance even when true heterogeneity is considerable. The confidence interval may be enormous. The random-effects weights, which depend on it, are then unstable, and the conventional confidence interval for the pooled effect is too narrow. Options include using an estimator with better small-sample properties, the Hartung-Knapp-Sidik-Jonkman adjustment for the pooled interval, and a Bayesian analysis with an informative or weakly informative prior for tau, taking the prior from empirical distributions of heterogeneity in meta-analyses of the same type. Whatever is done, the report should say plainly that the data contain little information about the between-study variance. See choosing priors in Bayesian meta-analysis.

Tau-squared for different effect measures

Tau squared is specific to the effect measure and the scale on which pooling is done. A value for log odds ratios cannot be compared with one for log risk ratios, or with one for standardized mean differences. For binary outcomes the heterogeneity on the odds ratio scale is usually somewhat different from that on the risk ratio scale, and tau squared depends on how common the outcome is. For this reason, comparing tau squared across meta-analyses is only meaningful for the same effect measure and a similar kind of outcome. Empirical studies of collections of meta-analyses give typical values for different outcome types and comparisons, which are used to interpret an observed value and as priors in Bayesian analyses. A tau squared of 0.1 on the log odds ratio scale, for instance, is regarded as moderate in many clinical settings, but this has to be judged in context.

Reporting

A report should name the estimator of tau squared, give the estimate of tau squared and of tau with a confidence interval where possible, state I squared and the Q statistic, and give a prediction interval. It should say how the estimator was chosen and whether results were tested with another. Software and version are stated, because defaults differ: some programs default to DL, others to REML. The interpretation should translate tau into practical terms. Where the number of studies is small, the report should state that the estimate is imprecise. PRISMA 2020 asks for the methods used to explore heterogeneity and the results. See PRISMA 2020.

Common mistakes

  • Using the DerSimonian-Laird estimator by default without noting its limits.
  • Reporting tau squared without its square root, which is hard to interpret.
  • Reading an estimate of zero as proof that there is no heterogeneity.
  • Comparing tau squared across different effect measures.
  • Ignoring the uncertainty of tau squared when the number of studies is small.
  • Not stating the estimator, so the result cannot be reproduced.

Support

Choice and justification of estimators, with sensitivity analyses, is part of the meta-analysis service, and Bayesian treatment of heterogeneity is covered by the Bayesian meta-analysis service.

Get a quoteSend your data and the questions you want answered.

Frequently asked questions

What is the difference between tau-squared and I-squared?

Tau squared is the variance of true effects, in the units of the effect measure squared. I squared is the share of total variability due to heterogeneity, as a percentage. Tau squared measures how large the differences are, and I squared how large a share they are.

Which estimator should I use?

Restricted maximum likelihood and Paule-Mandel are generally preferred to DerSimonian-Laird. State the choice, and check it as a sensitivity analysis if the conclusion could depend on it.

Can tau-squared be zero?

The estimate can be truncated at zero. That means the data show no excess variation, and does not prove that the true effects are identical.

Why is tau-squared hard to estimate?

Because it is based on differences between studies, and few studies give little information. Its confidence interval is typically wide.

How do I interpret tau?

As the standard deviation of true effects, in the units of the effect measure. For ratio measures, exp(tau) is the typical factor by which true effects differ from the average.

What can I do about tau-squared with only three studies?

Use an adjusted interval for the pooled estimate, consider a Bayesian analysis with a justified prior, and report that the data contain little information about heterogeneity.

References

  1. Veroniki AA, Jackson D, Viechtbauer W, et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Res Synth Methods. 2016;7(1):55-79.
  2. DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials. 1986;7(3):177-188.
  3. Paule RC, Mandel J. Consensus values and weighting factors. J Res Natl Bur Stand. 1982;87(5):377-385.
  4. Viechtbauer W. Bias and efficiency of meta-analytic variance estimators in the random-effects model. J Educ Behav Stat. 2005;30(3):261-293.
  5. Langan D, Higgins JPT, Jackson D, et al. A comparison of heterogeneity variance estimators in simulated random-effects meta-analyses. Res Synth Methods. 2019;10(1):83-98.
  6. Viechtbauer W. Confidence intervals for the amount of heterogeneity in meta-analysis. Stat Med. 2007;26(1):37-52.
  7. Turner RM, Davey J, Clarke MJ, Thompson SG, Higgins JPT. Predicting the extent of heterogeneity in meta-analysis, using empirical data from the Cochrane Database of Systematic Reviews. Int J Epidemiol. 2012;41(3):818-827.
  8. IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. 2014;14:25.
  9. Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.