What meta-regression is
Meta-regression is a regression analysis in which the outcome is the effect size from each study of a meta-analysis and the explanatory variables are characteristics of the studies. It asks whether the effect varies systematically with something that differs between studies, such as the dose of an intervention, the length of follow-up, the average age of participants, the year of publication, the risk of bias of the study or the setting in which it was done. A subgroup analysis is the special case in which the characteristic is a category. The method page gives the general outline.
The reason for doing it is that a single pooled effect is often not the interesting result. If studies disagree, readers want to know why, and whether the effect is larger in some circumstances than in others. Meta-regression offers a way of examining that, and of estimating how much of the between-study variation a characteristic accounts for.
It also has well-known weaknesses, which this page describes at length because they determine when the analysis is worth doing and how its results should be read. Many published meta-regressions would not pass the checks listed below.
When meta-regression is appropriate
Meta-regression is worth doing when four conditions are met.
- There is real heterogeneity to explain, or a hypothesis about an effect modifier that matters for practice or theory.
- The moderators were chosen in advance, on the basis of a rationale, and are few. Examining many characteristics and reporting the ones that reach significance produces spurious findings.
- There are enough studies. A commonly cited guide is at least ten studies for each characteristic modeled, and fewer studies make the analysis unreliable.
- The characteristic varies across studies and is reported consistently, so that it can be coded without much guesswork.
It is the wrong tool when studies are few, when the moderator has little variation, when it is confounded with another characteristic, or when the aim is to draw conclusions about individual participants, which study-level data cannot support. A feasibility check on the coded data establishes which situation applies before any model is fitted.
What the service includes
- Moderator selection. A short list of characteristics, chosen with a stated rationale and recorded in the analysis plan before results are seen.
- Coding of study characteristics with a codebook, so that the variables are defined consistently. See data extraction and coding.
- Feasibility check of the number of studies, the variation in each moderator, and correlations between moderators.
- Model fitting. Mixed-effects meta-regression with the estimator and test justified, including tests that account for the small numbers of studies.
- Interpretation of results, including the amount of heterogeneity explained and the limits of the inference.
- Sensitivity analyses and checks for influential studies.
- Figures and tables, including bubble plots and subgroup forest plots.
- Methods and results text and the code for reproduction.
How the models work
The standard approach is a mixed-effects meta-regression. It extends the random-effects model by allowing the mean effect to depend on the moderators and by estimating the residual between-study variance that remains after they are taken into account. Each study is weighted by the inverse of the sum of its sampling variance and this residual variance, so more precise studies count more, but less so than in a fixed-effect model.
The coefficient for a moderator is the estimated change in the effect size for each unit of the moderator, with a confidence interval and a test. A fixed-effect meta-regression, which assumes no residual heterogeneity, gives intervals that are too narrow when heterogeneity remains and is rarely appropriate. The conventional test for coefficients in a mixed-effects model also tends to reject too often when studies are few. An adjustment of the Knapp-Hartung type, which uses a t distribution and an adjusted variance, gives better control of the error rate and is widely recommended, and permutation tests can be used where even this is doubtful.
The proportion of heterogeneity explained, often expressed as an R² analogue, is reported with caution, because it is imprecisely estimated when studies are few and can be negative or truncated to zero. Software for these models includes the metafor and meta packages in R and the corresponding procedures in other statistical packages, and the version and settings are reported. See the guide on tau-squared for the between-study variance.
Pitfalls that need to be addressed
Six problems recur, and an analysis should address each of them explicitly.
- Too many moderators
- Testing many characteristics inflates the chance of a false positive. The number is limited, chosen in advance, and any exploratory analysis is labeled as such and adjusted or presented as hypothesis-generating.
- Too few studies
- With fewer than about ten studies per moderator, estimates are unstable and the analysis has little power to detect a real effect. Absence of a significant moderator is then not evidence of absence.
- Ecological bias
- An association between the average value of a characteristic across studies and the effect does not necessarily hold for individuals within studies. A drug may appear more effective in studies with older participants, for example, without being more effective in older participants. Participant-level data are needed to answer the individual question.
- Confounding between moderators
- Study characteristics are correlated. Studies with a higher dose may also be longer and more recent, so that a dose effect cannot be separated from a time effect. The correlations are examined and the interpretation is correspondingly modest.
- Observational nature
- Study characteristics are not randomized. A meta-regression is an observational analysis of studies, and its findings suggest hypotheses and do not prove causes.
- Data dredging
- Reporting only the models that gave interesting results hides the many that did not. The plan, the number of models fitted and all results are reported.
Subgroup analysis and its relation to meta-regression
A subgroup analysis compares the pooled effects across categories of a study characteristic, such as high versus low risk of bias or adult versus pediatric participants. It is easier to present and to understand than a regression, but it has the same weaknesses of limited power, ecological bias and observational comparison, and some of its own. The central rule is that the claim of a difference between subgroups rests on a formal test of interaction. It is not enough that the effect is significant in one subgroup and not in another, because a difference in significance is not a significant difference. Where a continuous moderator has been split into categories for convenience, information is lost, and a regression on the continuous variable is usually better. See subgroup analysis in meta-analysis.
Sensitivity analyses and influence
Because a meta-regression often rests on a modest number of studies, a single influential study can drive a result. Influence diagnostics, such as leave-one-out analysis and measures of each study's effect on the coefficients, show whether this is so. Sensitivity analyses also vary the estimator of the between-study variance, the test used for the coefficients, and the handling of studies with extreme values or substantial missing data on the moderator. Where results change, the report says so. The aim is to show how robust a finding is, not to find the version of the model in which it appears. See sensitivity analysis.
Reporting
The report gives the rationale for each moderator, the moment at which it was chosen, how it was coded, the number of studies and the range of values, the model and estimator, the test used, the coefficients with confidence intervals, the heterogeneity before and after adjustment, and the sensitivity results. A bubble plot, in which each study is plotted against the moderator with a symbol sized by its weight and the regression line overlaid, is the standard figure. Any analysis that was not prespecified is declared as post hoc. The wider review is reported against PRISMA 2020, and the data and code are made available so that others can check the analysis.
Deliverables
- Moderator plan with rationale and a codebook.
- Feasibility assessment of the number of studies, variation and collinearity.
- Results with coefficient estimates, intervals, tests and heterogeneity explained.
- Bubble plots and forest plots in publication-ready formats.
- Sensitivity and influence analyses.
- Methods and results text, with analysis code and output logs.
Get a quoteTell us your outcome, the number of studies and the moderators you have in mind.
Optional extension: full manuscript and submission
Full manuscript and submission package
When the scope includes manuscript preparation and submission, the package also contains the following.
Full manuscript draft
The rationale for each moderator, the results and bubble plots, with the limits of study-level inference stated.
Reporting checklist
PRISMA 2020, with analyses that were not prespecified labeled as post hoc.
Submission package
A journal recommendation, the manuscript formatted to that journal, a cover letter and supplementary files.
Revision round
Responses to queries about moderators, power and multiple testing.
Manuscript preparation follows the research integrity and authorship statement. The researchers who conceived the study and interpret its findings remain responsible for the content and its conclusions, and contributions that do not meet authorship criteria are acknowledged.
Limitations
Meta-regression cannot be better than the data allow. With few studies it is underpowered and its estimates are unstable. Associations are between studies and are exposed to ecological bias and confounding, so they generate hypotheses more often than they test them. Study characteristics reported poorly or inconsistently cannot be coded reliably, and missing moderator data reduce the number of studies that can be analyzed. The analysis is also easy to over-interpret: a significant coefficient in a modest sample is a weak basis for a strong claim. These limits belong in the report, and they are a reason to prefer a participant-level analysis if the question concerns individuals and the data can be obtained.
This service provides research and evidence-synthesis support. It does not provide clinical advice, and results are not patient-specific guidance.
How long does it take?
The time depends on how many moderators must be coded and checked, whether the characteristics are reported consistently in the included studies, and how many models and sensitivity analyses are planned. Coding the studies is usually the longer part, and the analysis itself is quicker. After the feasibility check, the schedule and its dependencies are agreed so that a fixed deadline can be planned for.
Frequently asked questions
How many studies do I need for a meta-regression?
A common guide is at least ten studies for each moderator modeled. With fewer, estimates are unstable and the analysis has little power, so a non-significant result should not be read as showing no effect.
What is the difference between subgroup analysis and meta-regression?
Subgroup analysis compares pooled effects across categories, and meta-regression relates effect sizes to categorical or continuous characteristics in a single model. Both are observational comparisons between studies, and both should rely on a formal test of interaction.
What is ecological bias?
It is the mistake of assuming that an association seen across study averages holds for individuals. A drug can look more effective in studies of older people without being more effective in older individuals. Participant-level data are needed to answer the individual question.
Can I test many moderators to find what explains heterogeneity?
Testing many moderators and reporting those that are significant gives misleading results. The moderators should be few and chosen in advance, and any exploratory analysis should be labeled as such.
Which software is used?
R packages such as metafor and meta, and equivalent procedures in other statistical software, are common. The software and settings are named in the methods and delivered with the code.
Does this service provide clinical advice?
No. It provides research and evidence-synthesis support. Results are not advice about the care of an individual patient.
References
- Thompson SG, Higgins JPT. How should meta-regression analyses be undertaken and interpreted? Stat Med. 2002;21(11):1559-1573.
- Thompson SG, Sharp SJ. Explaining heterogeneity in meta-analysis: a comparison of methods. Stat Med. 1999;18(20):2693-2708.
- Higgins JPT, Thompson SG. Controlling the risk of spurious findings from meta-regression. Stat Med. 2004;23(11):1663-1682.
- Knapp G, Hartung J. Improved tests for a random effects meta-regression with a single covariate. Stat Med. 2003;22(17):2693-2710.
- Berlin JA, Santanna J, Schmid CH, Szczech LA, Feldman HI. Individual patient- versus group-level data meta-regressions for the investigation of treatment effect modifiers: ecological bias rears its ugly head. Stat Med. 2002;21(3):371-387.
- Borenstein M, Higgins JPT. Meta-analysis and subgroups. Prev Sci. 2013;14(2):134-143.
- Lopez-Lopez JA, Marin-Martinez F, Sanchez-Meca J, Van den Noortgate W, Viechtbauer W. Estimation of the predictive power of the model in mixed-effects meta-regression: a simulation study. Br J Math Stat Psychol. 2014;67(1):30-48.
- Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1-48.
- Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71