Method

Meta-regression: the method

Meta-regression extends meta-analysis by relating the effect sizes of studies to characteristics of those studies. It can suggest why results differ, but it is easily misused. This page describes the models, how to read them, and the safeguards that make the analysis credible.

What meta-regression is

A conventional meta-analysis reports a single average effect, and an assessment of how much studies differ. Meta-regression goes a step further and asks whether the differences are associated with identifiable characteristics of the studies. It is a regression in which each study is one observation, the outcome is the study's effect size, and the explanatory variables, called moderators or covariates, are characteristics such as the average age of participants, the dose of the intervention, the duration of follow-up, the year of the study, the setting or the risk of bias. A coefficient for a moderator estimates how much the effect changes for each unit of the moderator.

The procedure is a natural extension of meta-analysis and uses the same weights, with similar assumptions. It is also an observational analysis of studies, not an experiment, which matters for how its results are understood. The service is described on the meta-analysis and meta-regression service pages, and this page covers the method.

The mixed-effects model

The standard model is mixed-effects meta-regression. For study i with estimate y_i and sampling variance v_i, the model says that y_i has a mean given by a linear function of the study's moderators, plus a study-specific random effect with variance tau squared, plus sampling error. The coefficients are estimated by weighted least squares, with weights equal to 1 divided by the sum of v_i and the residual between-study variance. The residual variance, estimated by the same methods as in random-effects meta-analysis, is the heterogeneity that remains after the moderators are accounted for. If the moderators explain much of the heterogeneity, this residual falls.

A fixed-effect meta-regression, which assumes that the moderators fully explain the variation, yields standard errors that are too small when residual heterogeneity remains, and it is rarely appropriate. Mixed-effects models are the default. Binary and categorical moderators are handled by indicator variables, and a model with a single categorical moderator is equivalent to a subgroup analysis in which the between-group variance is shared. Continuous moderators are best left continuous, since dividing them into categories loses information and can create artificial thresholds. Centering a continuous moderator, for instance at its mean, makes the intercept interpretable as the effect at an average value.

Tests and intervals

The coefficient of a moderator is tested with a z test in the conventional approach. With a small number of studies, this test rejects the null hypothesis too often, because it ignores the uncertainty in the estimate of the residual variance. The Knapp-Hartung adjustment replaces the normal distribution with a t distribution and modifies the variance estimate, and it has much better control of the false-positive rate. It is widely recommended, with a note that, as with the corresponding adjustment for the mean effect, the adjusted variance can in some circumstances be smaller than the unadjusted one. Permutation tests offer another route that does not depend on the distributional approximation and is particularly useful when many moderators are tested.

The omnibus test of whether any moderator is associated with the effect, and the tests for individual coefficients, are reported with their degrees of freedom. When several moderators are in the model, they are correlated, and the coefficient for each is its association with the effect holding the others constant, which needs enough studies to be estimable.

Variance explained

It is natural to ask how much of the heterogeneity a moderator explains. A common measure compares the residual between-study variance in the model with that in the model without moderators, and expresses the reduction as a percentage, an analogue of R squared. It should be read with caution. The estimate is highly imprecise with few studies and can be negative, in which case it is reported as zero, or it can be inflated. It is also a measure of explanation among the observed studies, not of prediction in new ones. A moderator that accounts for most of the heterogeneity among the included studies may not do so in future studies. The measure is reported with its uncertainty, if that can be estimated, or with a statement of its limits.

How many studies are needed

Meta-regression has little power unless studies are many. A rule of thumb often cited, including in the Cochrane Handbook, is that there should be at least ten studies for each moderator in the model, and fewer studies give unreliable results. In practice, power also depends on how much the moderator varies across studies, how large the effect of the moderator is, and how much residual heterogeneity there is. A moderator that takes nearly the same value in all studies cannot be evaluated. For this reason, an analysis that finds nothing with a modest number of studies is inconclusive and not evidence of no effect, and an analysis that finds something with few studies should be viewed skeptically, since such results often fail to replicate.

Ecological bias and confounding

Two threats to validity stand out. The first is ecological bias, also called aggregation bias. Meta-regression uses study-level averages, such as the mean age of participants, and relates them to the effect. An association across studies does not mean the same association exists within studies, between individuals. A treatment may appear more effective in studies that enrolled older patients, even if the effect does not vary with age among patients within any study, because other features of such studies, such as the setting or the definition of the outcome, change with average age. Only participant-level data, in an individual participant data meta-analysis, can answer questions about individuals reliably.

The second is confounding between moderators. Study characteristics are correlated: studies with higher doses may be longer, more recent and in different countries, so the effect of one cannot easily be separated from the others. A moderator that is associated with the effect may be a proxy for another that was not coded. The analyst examines the correlations among moderators, avoids entering highly collinear variables together, and states that the associations are exploratory. See individual participant data meta-analysis.

Choosing moderators and multiplicity

The analysis is credible when the moderators were chosen before the results were seen, for stated reasons, and are few. A moderator should be justified by a hypothesis or by theory, for example that the effect may depend on dose or on baseline severity, and not by what is available in the dataset. Testing many characteristics and reporting those that reached significance produces misleading results, as the chance of at least one false positive rises quickly with the number of tests. The remedies are to restrict the number of moderators, to specify them in the protocol, to report all that were tested, and to use methods that control the error rate, such as permutation tests that account for multiple testing. Analyses that were not pre-specified are labeled as exploratory and treated as generating hypotheses.

A numerical illustration

The table shows eight simulated studies in which the effect (a standardized mean difference) is plotted against the mean age of participants. The data are invented for illustration and do not come from real research. With weights equal to the inverse of each study's variance plus an assumed residual between-study variance of 0.005, a weighted regression of effect on mean age gives a slope of 0.011 per year, with a standard error of 0.005, so that the effect is estimated to rise by about 0.11 for each additional ten years of mean age. The ratio of the slope to its standard error is 2.02.

Simulated studies: effect against mean age
StudyMean ageEffect (SMD)Standard error
Study 1380.120.14
Study 2440.180.11
Study 3490.220.16
Study 4520.310.10
Study 5570.280.13
Study 6610.400.12
Study 7660.370.15
Study 8710.520.17

With only eight studies and one moderator the estimate is imprecise, and the result would be tested with the adjusted method described above before anything was claimed. The association is between study averages, so it does not show that older individuals respond more strongly, and other characteristics that rise with mean age in these studies could explain it. This is the kind of result that a meta-regression produces: a clear, interpretable slope and a set of reasons for caution.

Presenting the results

The standard figure is the bubble plot, in which each study is plotted at its moderator value against its effect size, with the size of the point proportional to its weight, and the regression line and its confidence band overlaid. It shows the data behind the model, reveals influential points and nonlinearity, and lets readers judge whether the association is driven by one or two studies. A table reports the coefficients, their intervals and tests, the residual heterogeneity and the number of studies. For categorical moderators, forest plots grouped by category show the subgroup estimates.

Sensitivity and influence

Because meta-regressions often rest on few studies, a single influential study can drive the result. Leave-one-out analyses and influence diagnostics show whether this is the case, and the analysis is repeated with different estimators of the residual variance and with the unadjusted and adjusted tests. Studies with missing moderator data are excluded, so the number of studies actually analyzed is reported, and if missingness may be related to the effect, this is considered. Where the results change with reasonable alternatives, the report says so.

Reporting

The report states the rationale and timing for each moderator, how it was coded, the number of studies and the range of values, the model and estimator, the test used, the coefficients with intervals, the residual heterogeneity and the variance explained with its caveats, and the sensitivity results. It declares which analyses were not specified in the protocol and notes the limits of inference from study-level data. The review follows PRISMA 2020, with the data and code shared so that others can examine the analysis.

How we can help

Support

Support for meta-regression

The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.

Get a quoteTell us your question, the number of studies and your target journal.

Frequently asked questions

What is the difference between subgroup analysis and meta-regression?

A subgroup analysis compares pooled effects across categories. A meta-regression relates effect sizes to categorical or continuous moderators in one model. Both are observational comparisons across studies.

How many studies do I need?

A common guide is at least ten studies per moderator, and fewer give unreliable results. A non-significant finding with few studies is inconclusive.

What is ecological bias?

The mistake of assuming that an association across study averages holds for individuals within studies. Participant-level data are needed for questions about individuals.

Should I use the Knapp-Hartung adjustment?

It is widely recommended for tests of coefficients because it controls the false-positive rate better when studies are few. Permutation tests are another option.

Can meta-regression prove that a moderator causes a difference in effect?

No. Study characteristics are not randomized and are correlated with each other, so meta-regression suggests hypotheses and cannot establish causes.

Why should moderators be chosen in advance?

Because testing many characteristics and reporting those that are significant produces false positives. Pre-specified moderators with a stated rationale give results that can be trusted more.

References

  1. Thompson SG, Higgins JPT. How should meta-regression analyses be undertaken and interpreted? Stat Med. 2002;21(11):1559-1573.
  2. Thompson SG, Sharp SJ. Explaining heterogeneity in meta-analysis: a comparison of methods. Stat Med. 1999;18(20):2693-2708.
  3. Higgins JPT, Thompson SG. Controlling the risk of spurious findings from meta-regression. Stat Med. 2004;23(11):1663-1682.
  4. Knapp G, Hartung J. Improved tests for a random effects meta-regression with a single covariate. Stat Med. 2003;22(17):2693-2710.
  5. Berlin JA, Santanna J, Schmid CH, Szczech LA, Feldman HI. Individual patient- versus group-level data meta-regressions for the investigation of treatment effect modifiers: ecological bias rears its ugly head. Stat Med. 2002;21(3):371-387.
  6. Borenstein M, Higgins JPT. Meta-analysis and subgroups. Prev Sci. 2013;14(2):134-143.
  7. Lopez-Lopez JA, Marin-Martinez F, Sanchez-Meca J, Van den Noortgate W, Viechtbauer W. Estimation of the predictive power of the model in mixed-effects meta-regression: a simulation study. Br J Math Stat Psychol. 2014;67(1):30-48.
  8. Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  9. Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1-48.
  10. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.