Guide

What is heterogeneity in meta-analysis?

Heterogeneity is variation in the true effects between studies in a meta-analysis, beyond what chance would produce. It is expected, informative and often large. This guide explains its types and measures, with a worked example, and what to do when it is present.

What heterogeneity is

When several studies address the same question, their results differ. Some of the difference is chance: even if every study were estimating exactly the same true effect, their estimates would scatter around it because each studied a different sample. Heterogeneity is the part of the difference that goes beyond this, because the true effects themselves differ. One trial may enroll sicker patients, another may deliver the treatment in a different way, a third may measure the outcome earlier. Such differences make the effect genuinely larger in some studies than in others.

Heterogeneity is not a flaw in a meta-analysis. It is a property of the evidence, and it carries information about how consistent an effect is and under what circumstances it varies. What matters is whether it is recognized, measured, explored and taken into account when interpreting the pooled result. A review that reports a single average from highly heterogeneous studies without comment gives a misleading impression of uniformity. The pooled estimate and the models that use it are described in meta-analysis.

Three kinds of heterogeneity

Clinical heterogeneity
Variation in participants, interventions or outcomes: differences in age, severity, setting, dose, delivery, comparator or the way an outcome is defined. It is a matter of subject-matter judgment and should be considered when the review is designed.
Methodological heterogeneity
Variation in study design and risk of bias: randomized and non-randomized designs, blinding, concealment, follow-up, handling of missing data and the adjustment used in observational studies.
Statistical heterogeneity
Variation in the observed effects beyond chance. It is the consequence of the first two, and of anything else that makes true effects differ, and it is what the statistics measure.

Clinical and methodological differences are the reasons; statistical heterogeneity is the evidence of them. The two are not the same, because studies can differ clinically and give similar results, and can appear statistically homogeneous when they are few or imprecise but still differ. The decision to pool should rest on the clinical and methodological comparability of the studies as well as on the statistics.

How heterogeneity is measured

Four quantities are used. The Q statistic, due to Cochran, is the weighted sum of squared differences between each study's estimate and the pooled estimate. If all studies shared one true effect, Q would follow a chi-squared distribution with one fewer degree of freedom than the number of studies, so a Q much larger than its degrees of freedom indicates heterogeneity. Tau squared is the estimated variance of the true effects, on the scale of the effect measure, and tau, its square root, is the standard deviation. I squared is the percentage of the total variability in estimates that is due to heterogeneity and not chance, calculated as the excess of Q over its degrees of freedom divided by Q. The prediction interval gives the range in which the true effect of a new study would be expected to lie.

Each has a role. Q tests whether heterogeneity exists, although with low power for few studies. Tau squared and tau quantify its size in the units of the effect. I squared puts it on a relative scale, but it depends on the precision of the studies. The prediction interval shows what the heterogeneity means for the effect in practice. They should be reported together. See the guides on I squared, tau squared and prediction intervals.

A worked example

Five simulated studies report standardized mean differences. They are invented for illustration and are not real research.

Simulated studies with substantial heterogeneity
StudyEffect (SMD)Standard error
Study A0.100.12
Study B0.550.15
Study C0.200.10
Study D0.750.18
Study E0.350.14

The inverse-variance (fixed-effect) pooled estimate is 0.31. The Q statistic is 12.89 on 4 degrees of freedom, far above its expected value of 4 under no heterogeneity. I squared is (12.89 minus 4) divided by 12.89, which is 69 percent. The DerSimonian-Laird estimate of tau squared is 0.0392, so tau is 0.20, meaning that true effects typically differ from their average by about 0.20 standard deviation units. Under random effects the pooled estimate is 0.36, with a 95 percent confidence interval from 0.15 to 0.57. The prediction interval, using a t distribution with 3 degrees of freedom, runs from -0.35 to 1.08, so a new study could plausibly find an effect anywhere from near zero to well above the average. The single average of 0.36 describes these studies poorly, and the interesting question is why Study A and Study D differ so much.

Why studies differ

The causes of heterogeneity are many, and listing candidate causes in advance is part of planning a review. Differences in the participants are the commonest: age, sex, disease severity, comorbidity, baseline risk and genetic or cultural background can all change how much a treatment helps. Differences in the intervention include dose, intensity, duration, delivery format, the skill of those delivering it and co-interventions. Differences in the comparator, for example placebo against usual care, change the contrast being estimated. Differences in outcome definition and timing, such as measuring pain at one week and at six months, matter. Differences in design and conduct, such as inadequate randomization or blinding, can inflate or deflate effects. And chance can create apparent heterogeneity when many subgroups are examined.

Some heterogeneity arises from the analysis. Different effect measures, such as odds ratios and risk ratios, can show different patterns of consistency, and studies with very different baseline risks may be homogeneous on one scale and heterogeneous on another. Errors in data extraction can create an outlier that looks like heterogeneity, so a check of extracted values is part of any investigation.

Exploring heterogeneity

When heterogeneity is present, the review should try to understand it. The usual tools are subgroup analysis, meta-regression and sensitivity analysis. Subgroup analysis compares pooled effects across categories of a study characteristic, using a test for interaction to decide whether the groups differ. Meta-regression relates effects to characteristics, including continuous ones, in a model. Sensitivity analyses, such as leaving out each study in turn, excluding studies at high risk of bias, or removing an outlier, show whether the heterogeneity depends on a few studies. A plot of the results, such as a forest plot ordered by a characteristic, a Galbraith plot or a leave-one-out plot, often shows the source at a glance.

These analyses should be planned, with a small number of characteristics chosen in advance for stated reasons, since examining many leads to false findings. They are observational comparisons across studies, so they suggest explanations and do not prove them, and with few studies they have little power. The explanations should be treated as hypotheses. See subgroup analysis.

Seeing heterogeneity in plots

Numbers describe heterogeneity, and plots show where it comes from. In a forest plot, heterogeneity appears as study intervals that do not overlap much, with squares scattered on both sides of the pooled diamond. A Galbraith or radial plot graphs each study's standardized estimate against its precision, with a band around the pooled line, and studies outside the band are the ones driving the excess variation. A leave-one-out plot shows the pooled estimate with each study omitted in turn, and a large shift on removing one study points to an influential outlier. A plot of effect size against a candidate moderator, with bubbles sized by weight, shows whether the variation tracks a characteristic. None of these is a test, but each helps decide which explanations are worth examining and communicates the pattern to readers more clearly than a single statistic.

Heterogeneity with few studies

With only a handful of studies, every measure of heterogeneity is imprecise. The Q test can miss real variation, I squared can swing between zero and a high value on small changes in the data, and tau squared has a very wide confidence interval, so that an estimate of zero is compatible with a large amount of heterogeneity. Reporting the uncertainty, for example the interval for tau squared and I squared, is more honest than reporting the point estimates alone. In practice with few studies, the sensible course is to reason about heterogeneity from the clinical and methodological differences between the studies, to use a random-effects analysis with an adjustment for the small number, to avoid claims about subgroups, and to say that the data contain little information about how much effects vary. A Bayesian analysis with a justified prior for heterogeneity is a further option.

Deciding whether to pool

The question is whether the studies are sufficiently similar that an average is meaningful. There is no cut-off in any statistic that answers it. A high I squared does not forbid pooling, because the studies may differ in size but not in direction, and a low I squared does not guarantee comparability, since a few imprecise studies may mask real differences. The decision rests on judgment about the question being asked. If the studies address different questions, because they involve different interventions, populations or outcomes, pooling is inappropriate whatever the statistics say. If they address the same question and differ in magnitude, a random-effects analysis with exploration and a prediction interval is appropriate. If the studies disagree in direction, so that some show benefit and some harm, a pooled average may hide an important finding, and the review should describe the pattern and consider not pooling.

Reporting and interpreting

A review should say how much heterogeneity was found, with Q, tau squared, I squared and the prediction interval, in a form that a reader can understand. It should say what was done about it: whether a random-effects model was used and why, which sources were explored and what was found. The interpretation should take heterogeneity into account: a statement that the treatment reduces the outcome on average, with substantial unexplained variation in the size of the effect, is more honest than one that omits the variation. In GRADE, unexplained inconsistency lowers the certainty of evidence. See GRADE.

Common mistakes

  • Using fixed thresholds for I squared as rules for what to do.
  • Choosing between fixed and random-effects models by a significance test of Q.
  • Reporting I squared without tau squared or a prediction interval.
  • Exploring many study characteristics and reporting the ones that reached significance.
  • Treating unexplained heterogeneity as unimportant because the pooled result is significant.
  • Dropping an outlier without a stated reason, instead of investigating it.

Support

Assessment and exploration of heterogeneity are part of the meta-analysis service and the meta-regression service.

Get a quoteSend your data and the questions you want answered.

Frequently asked questions

What is the difference between heterogeneity and inconsistency?

They are close in meaning. Heterogeneity is variation in true effects between studies. Inconsistency is the term used in GRADE and in network meta-analysis for unexplained variation in results, and for disagreement between direct and indirect evidence.

Is a high I-squared a reason not to pool?

Not by itself. It depends on whether the studies address the same question and whether they agree in direction. Explore the heterogeneity and report a prediction interval.

Why is the Q test unreliable?

It has low power when studies are few or small, so it can miss real heterogeneity, and high power when studies are many or large, so it can flag trivial differences.

What does a random-effects model do about heterogeneity?

It allows true effects to vary and estimates their average and spread. It accommodates heterogeneity in the analysis but does not explain it.

How should I explore heterogeneity?

With a few characteristics chosen in advance, using subgroup analysis or meta-regression, and with sensitivity analyses. Treat the results as hypotheses.

Can heterogeneity be zero?

It can be estimated as zero, especially with few studies, but that does not prove the true effects are identical, only that the data show no excess variation.

References

  1. Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med. 2002;21(11):1539-1558.
  2. Higgins JPT, Thompson SG, Deeks JJ, Altman DG. Measuring inconsistency in meta-analyses. BMJ. 2003;327(7414):557-560.
  3. Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.
  4. Rucker G, Schwarzer G, Carpenter JR, Schumacher M. Undue reliance on I-squared in assessing heterogeneity may mislead. BMC Med Res Methodol. 2008;8:79.
  5. Hoaglin DC. Misunderstandings about Q and 'Cochran's Q test' in meta-analysis. Stat Med. 2016;35(4):485-495.
  6. Cochran WG. The combination of estimates from different experiments. Biometrics. 1954;10(1):101-129.
  7. IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6(7):e010247.
  8. Thompson SG, Higgins JPT. How should meta-regression analyses be undertaken and interpreted? Stat Med. 2002;21(11):1559-1573.
  9. Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.