What subgroup analysis is
A subgroup analysis in meta-analysis splits the included studies into groups according to a characteristic, estimates a pooled effect within each group, and compares the groups. The characteristic may concern the studies, such as the dose of the intervention, the setting or the risk of bias, or it may concern the participants as summarized in each study, such as the age group or the sex of the participants, where studies report results separately for them. The reason for doing it is to see whether the effect varies in a systematic way, which is a way of explaining heterogeneity and of identifying where an intervention works better or worse.
The idea is simple and so is the trap. The comparison is easy to make and easy to over-interpret. Many published subgroup claims have not held up in later work. The analysis is observational, uses few studies in each group, and tests many possible characteristics. The safeguards described below address that, and the closely related method of meta-regression, which handles continuous characteristics, shares the same cautions.
The right test: a test for subgroup differences
The central error is to conclude that subgroups differ because the pooled effect is significant in one and not in the other. A difference in significance is not a significant difference. A subgroup can have a non-significant result because it has few or imprecise studies, even when its estimated effect is as large as that of the significant subgroup. The correct approach is a direct test of whether the effects in the subgroups differ, called a test for subgroup differences or a test of interaction.
The idea is easy to see in an example of two subgroups, which are invented for illustration. Studies 1 to 4 used a high dose, and Studies 5 to 8 a low dose. Pooling within each subgroup with fixed-effect weights gives 0.49 for the high-dose group (95 percent confidence interval 0.35 to 0.63) and 0.16 for the low-dose group (0.02 to 0.29). The difference between the two estimates is 0.33, with a standard error equal to the square root of the sum of the two squared standard errors, 0.099. The ratio, 3.37, is a z statistic, and its square, 11.4, is the between-subgroup heterogeneity statistic with one degree of freedom, equivalent to a chi-squared test. A z of 3.37 exceeds 1.96, so the data suggest a difference between the doses at the 5 percent level. For more than two subgroups, the test uses the between-group Q statistic with degrees of freedom one fewer than the number of subgroups.
Fixed or random effects within subgroups
Within the subgroups, the pooling can be fixed-effect or random-effects. With random effects, a further choice arises: whether to allow a separate between-study variance in each subgroup or to assume a common one. Separate variances are more flexible but are estimated from few studies and so are very imprecise, and a common variance borrows strength across subgroups but assumes the heterogeneity is the same. Many software packages use separate variances by default. With small subgroups, a common variance is often more stable, but its assumption should be considered. The test for subgroup differences, in either case, uses the pooled estimates and their standard errors from the chosen model, and with random effects the adjusted intervals described for small numbers of studies are also an option. The model should be specified in the protocol and reported.
Power is usually low
Subgroup analyses in meta-analysis typically have low power to detect differences between subgroups, for the same reason that single studies have low power for subgroups: dividing the data reduces precision. The power depends on the number of studies in each subgroup, their precision, the size of the true difference, and the residual heterogeneity. A difference of a size that would matter in practice may be undetectable with a few studies per group. The consequence is that a non-significant test for subgroup differences does not show that the effect is the same in the groups. It shows that the data do not distinguish them. Authors should say so, and should avoid concluding from a non-significant test that the intervention works equally in all groups, or that the pooled effect applies to every subgroup.
Low power has a second effect: when a significant difference is found with few studies, the estimated size of the difference is often inflated, which is why such findings frequently fail to replicate. See how many studies are needed.
Multiplicity and prespecification
Testing many characteristics leads to false positives. If twenty independent subgroup analyses are done, one is expected to be significant at the 5 percent level by chance. The risk is highest when the analyses are chosen after looking at the data. The safeguards are standard. A small number of subgroup analyses is specified in the protocol, each with a rationale from theory, mechanism or earlier evidence, and with a prediction of the direction of the difference. Analyses that were not prespecified are labeled post hoc and treated as generating hypotheses. All analyses that were run are reported, including those without a finding, and where many are run, an adjustment for multiple testing or a more cautious interpretation is applied. A subgroup finding that is prespecified, in the predicted direction, supported by a plausible mechanism and consistent across related analyses is credible. One that has none of these is much less so.
Study-level characteristics and ecological bias
When the characteristic is a property of the study, such as the dose or the setting, the subgroup comparison is a comparison between studies. It is vulnerable to confounding: studies with high doses may also differ in population, duration or risk of bias, and the difference observed may be due to those. Where the characteristic describes participants, such as the average age of the sample, a comparison across studies with older and younger samples is subject to ecological bias, because the association across study averages may differ from the association within studies. A finding that studies with older participants show larger effects does not mean older individuals respond more. The reliable approach is a comparison within trials, available from the trials' own subgroup results and best pooled with the interaction estimated within each trial, or from individual participant data. See individual participant data meta-analysis.
Continuous characteristics and meta-regression
Splitting a continuous characteristic, such as dose or age, into categories loses information, creates arbitrary cut-points, and invites choice of the cut-point that gives the most interesting result. When the characteristic is continuous, a meta-regression of effect on the characteristic is usually better. It uses the full range of values, gives an estimate of the change in effect per unit, and can include several characteristics. The same cautions apply, including the need for adequate numbers of studies, usually at least ten per characteristic. A subgroup analysis with a natural categorical variable, such as study design or the type of comparator, remains appropriate. See meta-regression.
Presenting a subgroup analysis
The usual figure is a forest plot grouped by subgroup, with a pooled estimate for each group, and, if useful, an overall estimate. The test for subgroup differences, with its statistic, degrees of freedom and p-value, is printed beneath. Heterogeneity statistics within each subgroup and the number of studies in each are shown. The text states which subgroup analyses were prespecified, gives the rationale and the direction predicted, and interprets the result in line with the cautions: a finding is described as suggestive, and its consistency with other evidence is discussed. The report should also say if the overall conclusion changes depending on the subgroup. See how to read a forest plot.
Judging whether a subgroup effect is believable
Several criteria help to judge whether an apparent subgroup difference is credible. Was the characteristic specified before the analysis, with a hypothesis about the direction? Was the comparison made within studies or only between them? Is the difference statistically significant on a test of interaction, and is it large enough to matter in practice? Was it one of a small number of analyses, or one of many? Is it consistent across studies, or driven by one? Is there external support, such as a biologically or theoretically plausible mechanism, or evidence from other studies? A difference that satisfies most of these deserves attention, and one that satisfies few should be presented as a hypothesis. These criteria were developed for subgroups in trials and apply with equal force to meta-analysis, where the number of studies is often small. The assessment can be recorded in the report in a few sentences, and it helps readers to see the claim at its proper strength.
Reporting subgroup analyses
The methods section should list the subgroup analyses and say which were planned, give the reason for each, describe how subgroups were defined, state the model and the handling of heterogeneity within groups, and name the test for differences. The results should give the pooled estimate and interval for each subgroup, the number of studies and participants, the heterogeneity statistics and the test result. The discussion should interpret the differences with the cautions above, and should state the limitations, particularly low power and the observational nature of the comparison. PRISMA 2020 asks authors to report the methods for exploring heterogeneity and the results of those explorations. Where the analyses were not prespecified, the manuscript says so clearly. This level of reporting also helps in peer review, since reviewers of meta-analyses frequently ask about the basis for subgroup claims. See PRISMA 2020.
Common mistakes
- Comparing significance levels within subgroups instead of testing the difference between them.
- Running many subgroup analyses and reporting the significant ones.
- Choosing cut-points after seeing the data to produce a difference.
- Concluding that subgroups do not differ from a non-significant test with low power.
- Treating a between-study comparison as a within-person effect (ecological bias).
- Splitting a continuous variable when meta-regression would be better.
- Using separate heterogeneity variances in very small subgroups without noting the instability.
Support
Planning and running subgroup analyses and the related meta-regression is part of the meta-regression service and the meta-analysis service.
Get a quoteTell us your outcome, the number of studies and the characteristics of interest.
Frequently asked questions
How do I test whether subgroups differ?
Use a test for subgroup differences, an interaction test, which compares the pooled estimates between groups, not the significance of each group on its own.
Why is it wrong to say the effect was significant in one group and not the other?
Because a difference in significance is not a significant difference. A group can be non-significant simply because it has fewer or more imprecise studies.
How many subgroup analyses can I do?
A small number, specified in the protocol with a rationale. More analyses increase the risk of false positives.
What if the test for subgroup differences is not significant?
It means the data do not distinguish the groups, which can reflect low power. It does not show the effect is the same.
Should I categorize a continuous characteristic?
Usually not. A meta-regression on the continuous variable uses more information and avoids arbitrary cut-points.
Can subgroup analysis explain heterogeneity?
It can suggest explanations, but with few studies the explanations are tentative and should be treated as hypotheses.
References
- Borenstein M, Higgins JPT. Meta-analysis and subgroups. Prev Sci. 2013;14(2):134-143.
- Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
- Altman DG, Bland JM. Interaction revisited: the difference between two estimates. BMJ. 2003;326(7382):219.
- Thompson SG, Higgins JPT. How should meta-regression analyses be undertaken and interpreted? Stat Med. 2002;21(11):1559-1573.
- Higgins JPT, Thompson SG. Controlling the risk of spurious findings from meta-regression. Stat Med. 2004;23(11):1663-1682.
- Oxman AD, Guyatt GH. A consumer's guide to subgroup analyses. Ann Intern Med. 1992;116(1):78-84.
- Sun X, Briel M, Walter SD, Guyatt GH. Is a subgroup effect believable? Updating criteria to evaluate the credibility of subgroup analyses. BMJ. 2010;340:c117.
- Berlin JA, Santanna J, Schmid CH, Szczech LA, Feldman HI. Individual patient- versus group-level data meta-regressions for the investigation of treatment effect modifiers: ecological bias rears its ugly head. Stat Med. 2002;21(3):371-387.