Guide

What is I-squared?

I-squared is the most widely reported measure of heterogeneity in meta-analysis, and one of the most misread. It tells you what share of the variation between study results is due to real differences and not chance. It does not tell you how large those differences are. This guide explains what it measures, what it does not, and what to report with it.

What I-squared measures

In any meta-analysis, study estimates differ. Some of the difference is chance, because each study is a different sample, and some may be real, because the true effect varies. I squared, introduced by Higgins and Thompson in 2002, expresses the share of the total variation among estimates that is due to the real differences. An I squared of 0 percent means that the observed variation is no more than chance alone would produce. An I squared of 75 percent means that three quarters of the variation reflects genuine differences between studies and one quarter is chance. It was designed to be easy to interpret and independent of the number of studies, and it quickly became a standard item in meta-analysis reports.

It is computed from the Q statistic, the weighted sum of squared differences between each study estimate and the pooled estimate, and the degrees of freedom, which are the number of studies minus one. If Q is larger than its degrees of freedom, there is more variation than chance explains, and I squared is the excess as a share of Q. If Q is less than its degrees of freedom, I squared is set to zero. The general background is in the guide to heterogeneity.

How it is calculated

The formula is I squared equals Q minus df, divided by Q, times 100 percent, with a floor of zero. A related quantity, H, is the square root of Q divided by df, and I squared can be written as H squared minus one over H squared. H of 1 means no heterogeneity, and larger values mean more. The calculation uses only Q and the number of studies, so it does not depend on the effect measure, and can be computed for any meta-analysis.

A worked example uses 7 simulated studies, which are invented for illustration. The fixed-effect pooled estimate is 0.27, and Q is 9.71 on 6 degrees of freedom. I squared is (9.71 minus 6) divided by 9.71, which is 38 percent, and H is 1.27. A 95 percent confidence interval for H, calculated on the log scale with the standard error given by Higgins and Thompson, runs from 0.83 to 1.96, which translates into an interval for I squared of 0 to 74 percent. The width of that interval is the first lesson: with 7 studies, the data are compatible with heterogeneity ranging from small to very large, and the single figure conceals this.

Why I-squared is not an absolute measure

The most important limitation is that I squared depends on the precision of the studies, not only on how much true effects differ. Heterogeneity in the true effects is measured by tau squared, the between-study variance. I squared compares tau squared with the typical within-study variance. If the studies are small, with large within-study variances, a given tau squared is dwarfed by sampling error, and I squared is low. If the studies are large, with small variances, the same tau squared stands out, and I squared is high. The relationship is roughly I squared equals tau squared divided by the sum of tau squared and the typical within-study variance.

An example shows the effect. Take a between-study variance of 0.04, which means that true effects typically differ by about 0.2 units. If the studies have standard errors around 0.20, the typical within-study variance is 0.0400, and I squared is about 50 percent. If the studies have standard errors around 0.05, the within-study variance is 0.0025, and I squared is about 94 percent. The true variation in effects is the same in both cases, and I squared differs by more than 44 percentage points. The point was made in a detailed critique by Borenstein and colleagues, who recommended not treating I squared as an absolute measure of heterogeneity. The figures are illustrative.

The rough thresholds, and why not to rely on them

The Cochrane Handbook offers a rough guide to interpretation: 0 to 40 percent might not be important, 30 to 60 percent may represent moderate heterogeneity, 50 to 90 percent may represent substantial heterogeneity, and 75 to 100 percent considerable heterogeneity. The overlapping ranges are deliberate, and the Handbook adds that the importance of the observed value depends on the size and direction of effects and on the strength of the evidence for heterogeneity, as shown by the p-value of the Q test or the confidence interval for I squared. In practice the thresholds are often applied as if they were rules, with a result of 50 percent triggering a switch of model or a decision not to pool. Such use is not supported. The same value can mean different things for different effect measures and study sizes, and a high value from large, consistent studies, whose effects differ trivially, is very different from the same value from studies whose effects differ greatly in practical terms.

Uncertainty in I-squared

As the worked example shows, I squared is estimated imprecisely, especially when the number of studies is small. Its sampling distribution is wide, so a point estimate of zero is compatible with considerable heterogeneity, and a high estimate is compatible with moderate heterogeneity. Higgins and Thompson gave methods for confidence intervals, based on H, and they are available in meta-analysis software. Reporting the interval, or at least noting the imprecision, is advisable, and some journals require it. More refined intervals, based on the profile likelihood or the generalized Q statistic, perform better in some situations. When a meta-analysis has fewer than about ten studies, the interval is so wide that the statistic carries little information, and the conclusions about heterogeneity should rest on judgment about the studies themselves.

Where the statistic came from

Before I squared, the standard way to assess heterogeneity was the Q test, which asks whether the variation among studies is greater than chance, and gives a p-value. Its weaknesses were known: it has low power when studies are few or small, so it can miss heterogeneity that matters, and excessive power when studies are many or large, so it flags trivial differences. Users needed a measure of the extent of heterogeneity and not just a test. Higgins and Thompson proposed three descriptive statistics, H, R and I squared, and showed how they relate to Q. I squared won because it was intuitive, a percentage, and was claimed to be comparable across meta-analyses of different sizes and types. Later work has shown that the last claim is only partly true, because of the dependence on precision, but the statistic has remained, and it is built into every major meta-analysis program. Its role is best seen as a first descriptive summary and a prompt to look further.

I-squared, tau-squared and the prediction interval compared

Measures of heterogeneity compared
MeasureWhat it tells youScaleMain weakness
Q testWhether variation exceeds chanceChi-squared statistic and p-valueLow power with few studies; high power with many
I-squaredThe share of variation due to heterogeneityPercentage, 0 to 100Depends on study precision; imprecise with few studies
Tau-squared / tauThe size of the variation in true effectsEffect-measure unitsImprecise with few studies; depends on the estimator
Prediction intervalThe range of effects expected in a new studyEffect-measure unitsWide and unreliable with few studies

No single measure is enough. A report that gives all four, with a sentence of interpretation, serves the reader better than any one alone.

What to report alongside it

  • Tau-squared and tau. The between-study variance and standard deviation, on the scale of the effect measure, which tell the reader how large the real differences are. See tau-squared.
  • A prediction interval. The range of effects expected in a new study, which turns heterogeneity into a statement about what to expect. See prediction intervals.
  • The Q statistic, with its degrees of freedom and p-value, noting its low power with few studies and high power with many.
  • A confidence interval for I squared, or a note about its imprecision.
  • A plot, such as the forest plot or a Galbraith plot, showing the variation directly.

A short interpretation should follow, in words: how the studies differ, whether the differences are important in practice, and what was done about them.

Common misuses

  • Choosing fixed or random effects by I squared. The choice rests on the question and the similarity of the studies, as described in fixed-effect versus random-effects models.
  • Using a threshold to forbid pooling. A high value is a reason to explore, not an automatic bar.
  • Reading a low value as proof of homogeneity, when the studies are small and the statistic is imprecise.
  • Comparing I squared across meta-analyses with very different study sizes or effect measures, as if it were on one scale.
  • Reporting it without a number of studies, which makes it impossible to judge its reliability.
  • Using it for subgroup comparisons without a test of the difference between subgroups.

Related measures

Several related statistics appear in the literature. H, the square root of Q over its degrees of freedom, is the basis of the confidence interval. R, a ratio of standard errors under random and fixed-effect models, is another summary, rarely reported. The index D squared, the diversity measure, is similar to I squared but based on the change in variance between models. For network meta-analysis, versions of I squared are reported for the whole network and for comparisons. For diagnostic accuracy meta-analysis, I squared is not appropriate, because there are two correlated outcomes, and measures from the bivariate model are used instead. For dose-response and multivariate settings, other measures apply. Whenever a variant is used, the report states which one and how it was calculated. See diagnostic accuracy meta-analysis.

Support

Heterogeneity assessment and its interpretation are part of the meta-analysis service.

Get a quoteSend your data and the questions you want answered.

Frequently asked questions

What does an I-squared of 50 percent mean?

That about half of the total variability in the estimates is due to real differences between studies and half to chance. It does not say how large the real differences are.

Is a high I-squared bad?

Not necessarily. It may reflect precise studies with small differences in effect. It signals that heterogeneity should be explored and interpreted, not that the analysis is invalid.

Can I-squared be negative?

The formula can give a negative value when Q is smaller than its degrees of freedom, and it is set to zero in that case.

Should I report I-squared or tau-squared?

Both, since they answer different questions, together with a prediction interval. Tau-squared gives the size of the variation and I-squared its share.

Why does I-squared change when I add large studies?

Because it compares between-study variance with within-study variance. Larger studies reduce the within-study variance, so the same true variation makes up a bigger share.

How many studies are needed for a meaningful I-squared?

There is no minimum, but with fewer than about ten studies its confidence interval is wide, and it should be interpreted with caution.

References

  1. Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med. 2002;21(11):1539-1558.
  2. Higgins JPT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. BMJ. 2003;327(7414):557-560.
  3. Borenstein M, Higgins JPT, Hedges LV, Rothstein HR. Basics of meta-analysis: I-squared is not an absolute measure of heterogeneity. Res Synth Methods. 2017;8(1):5-18.
  4. Rucker G, Schwarzer G, Carpenter JR, Schumacher M. Undue reliance on I-squared in assessing heterogeneity may mislead. BMC Med Res Methodol. 2008;8:79.
  5. von Hippel PT. The heterogeneity statistic I-squared can be biased in small meta-analyses. BMC Med Res Methodol. 2015;15:35.
  6. Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  7. IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6(7):e010247.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.