The short answer
There is no minimum number of studies required to carry out a meta-analysis. Two studies that measure the same outcome in comparable ways can be combined, and Cochrane reviews have reported meta-analyses of two. The harder question is whether the combined estimate is worth having. A pooled result from two small, dissimilar studies may be less informative than reading them side by side, and it can give a false sense of precision.
The right question is therefore not how many studies are enough, but what conclusions the available studies can support. That depends on four things: how many participants and events the studies contain, how consistent their results are, how much they differ in design and population, and what the review intends to do beyond estimating one average effect.
When to pool and when not to
Pooling is appropriate when studies address the same question closely enough that an average is meaningful. Judging that is a matter of the populations, interventions, comparators and outcomes, which should be considered before looking at the results. If studies are too different, a pooled number describes nothing in particular, and a structured narrative synthesis, perhaps with a forest plot but no summary diamond, is more honest.
The reverse error is also real. Many reviews stop at a narrative when a meta-analysis would have been feasible, because the authors were unsure of the rules. In a review where several trials measure the same outcome in similar ways, pooling gives a more precise estimate and allows heterogeneity to be examined, and the protocol should say so. Where meta-analysis is not done, the reasons should be reported, and the SWiM guideline offers a structure for reporting synthesis without meta-analysis. A useful habit is to write down, before extraction, the criteria that will decide whether studies are similar enough to pool: the same outcome construct, a compatible time point, a comparable comparator and a similar enough population. Applying written criteria is easier to defend than a judgment made once the forest plot is on screen, and it keeps the decision to pool separate from whether the pooled result looks convenient.
Precision and power with few studies
A meta-analysis is usually judged by how precisely it estimates the effect and by its power to detect an effect of a given size. Both depend on the information in the studies and on heterogeneity. The table gives a simple illustration, for studies that are all the same size, with a within-study variance of 0.04 for the effect estimate, a true effect of 0.20 and, for the random-effects column, a between-study variance (tau-squared) of 0.04.
| Studies | Half-width, fixed | Half-width, random | Power, fixed | Power, random |
|---|---|---|---|---|
| 2 | 0.28 | 0.39 | 29% | 17% |
| 3 | 0.23 | 0.32 | 41% | 23% |
| 5 | 0.18 | 0.25 | 61% | 35% |
| 10 | 0.12 | 0.18 | 89% | 61% |
| 20 | 0.09 | 0.12 | 99% | 89% |
| 40 | 0.06 | 0.09 | 100% | 99% |
Three things follow. First, the interval shrinks with the square root of the number of studies, so each additional study helps less than the one before. Second, heterogeneity costs information: with the same number of studies, the random-effects interval is wider and power is lower. Third, power depends strongly on the setting. With these assumptions, about 8 studies would give 80 percent power for a fixed-effect analysis, and about 16 for the random-effects case. With a different within-study variance or amount of heterogeneity, the numbers change a great deal, and so these figures are not a rule. They show the way to think about the question.
Power in meta-analysis has been discussed by Jackson and Turner, who showed that it can be calculated in advance from assumptions about study size and heterogeneity, and that random-effects meta-analyses often have lower power than people expect. Power calculations are not usually reported, but looking at the expected width of the interval is a good way to judge whether a synthesis is worth doing.
Estimating heterogeneity with few studies
With a random-effects model, the between-study variance must be estimated from the variation among the studies, and with only a handful of studies that estimate is very uncertain. It can be zero when the true value is not, which makes the random-effects result collapse to the fixed-effect one, and it can be large and unstable. The confidence interval for the pooled effect from the usual method then tends to be too narrow, because it treats the estimated variance as if it were known.
The Hartung-Knapp-Sidik-Jonkman adjustment widens the interval to account for this uncertainty and has better coverage with few studies, although it can sometimes give a very wide interval, and an adjusted variance is occasionally narrower than the conventional one. A prediction interval needs at least three studies and is very wide with few. Bayesian methods with a weakly informative prior for the between-study standard deviation are another option, provided the sensitivity to the prior is examined. In every case, the right response to a small number of studies is to show the uncertainty clearly and not to conceal it behind a choice of model.
Subgroups, meta-regression and publication bias
Analyses that go beyond the overall effect need more studies. The Cochrane Handbook suggests that meta-regression should generally be considered only when there are at least ten studies for each characteristic modelled. With fewer, the association can be driven by one or two studies, the risk of false positives is high, and confounding between characteristics is hard to untangle. Subgroup analyses have similar limits, and a subgroup with two studies carries almost no information on its own. Both should be limited in number and prespecified.
Tests for funnel plot asymmetry are also unreliable with fewer than about ten studies, because power is low. A small review cannot say much about publication bias from its own data, and so the search and the prior knowledge of unpublished studies matter more.
What to do when there are few studies
Few studies are common, especially in new fields and in narrow questions. A review is not worthless because it finds little, and an honest account of a thin evidence base is useful to readers and funders. The following choices help.
- Say clearly how many studies and participants contributed to each analysis, and label results as based on few studies.
- Choose methods suited to few studies, such as a Hartung-Knapp adjustment, exact methods for sparse data, or a Bayesian model with a sensitivity analysis on the prior.
- Do not over-interpret subgroups and heterogeneity statistics, which are unstable.
- Rate the certainty of evidence frankly, since imprecision and inconsistency are domains in GRADE that few studies often trigger.
- Describe research gaps in a way that tells future researchers what study would add the most information.
- Consider whether the review should be a scoping or evidence-map review if the aim is to describe the extent of the literature and not to estimate an effect.
- Plan a living review if new studies are expected soon, so that the synthesis can be updated.
When there are many studies
A large number of studies brings its own difficulties. Heterogeneity can increase, and the meta-analysis may mix such different contexts that the average is hard to interpret, which makes prediction intervals and carefully planned subgroup analyses more important. Dependence between effect sizes becomes more common, because large reviews include multiple outcomes, arms and time points from the same sample. Screening and extraction take much longer and need more reviewers, so workload and error control should be planned. A very large evidence base also raises the question of whether a new review adds anything to existing ones, and the possibility of an umbrella review, which summarizes systematic reviews, should be considered.
Planning before you start
The number of eligible studies cannot be known before searching, but it can be estimated. A scoping search of one or two databases, or a look at existing reviews and registries, gives a sense of the likely volume. The result affects the choice of design, the team size, the timeline and the planned analyses. If the preliminary search suggests very few studies, the protocol can say in advance what will be done: which analyses will be run if there are at least a given number of studies, and what will be reported if there are not. This makes the plan robust to the outcome of the search and avoids decisions made after seeing the results.
Participants and events matter more than the count
Counting studies is a crude measure of information. Ten tiny trials may contain fewer participants than one large one, and for binary outcomes the number of events is what carries the information. A meta-analysis of five studies with several hundred events each can be more informative than one of twenty studies with a handful of events each. For this reason, GRADE judges imprecision partly by the total number of events or participants compared with the size of a single adequately powered trial, called the optimal information size, in addition to the width of the confidence interval.
The same logic underlies trial sequential analysis, which asks whether the accumulated information in a meta-analysis has reached the amount that would be needed to draw a firm conclusion, and adjusts the significance threshold for repeated looks at accumulating data. It is a useful tool for judging whether a significant pooled result in a small evidence base is likely to hold. Its assumptions need to be stated, and it is a complement to, not a replacement for, the assessment of risk of bias and heterogeneity.
How we can help
We can run a scoping search to estimate the number of eligible studies, advise on whether pooling is sensible, calculate expected precision and power, choose methods suited to few studies, and report limited evidence clearly. [OWNER VERIFICATION REQUIRED] The relevant services are meta-analysis, systematic review and statistical analysis.
Frequently asked questions
What is the minimum number of studies for a meta-analysis?
Technically two. Whether two studies give a useful result depends on how similar and how precise they are.
How many studies do I need for a funnel plot or Egger's test?
About ten or more, and of varying size. With fewer, the tests have low power.
How many studies do I need for meta-regression?
The Cochrane Handbook suggests at least about ten studies for each characteristic examined.
Should I use random effects with only a few studies?
You may, but the between-study variance is poorly estimated. Use the Hartung-Knapp adjustment or a Bayesian approach and report the uncertainty.
What if I find only two or three studies?
Report them transparently, pool only if they are comparable, avoid subgroup and publication bias analyses, and describe the certainty of evidence honestly.
Can a review be useful with no meta-analysis?
Yes. A structured narrative synthesis reported according to the SWiM guideline, or a scoping review, can be useful when pooling is not appropriate.
References
- Jackson D, Turner R. Power analysis for random-effects meta-analysis. Res Synth Methods. 2017;8(3):290-302.
- Valentine JC, Pigott TD, Rothstein HR. How many studies do you need? A primer on statistical power for meta-analysis. J Educ Behav Stat. 2010;35(2):215-247.
- IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. 2014;14:25.
- Campbell M, McKenzie JE, Sowden A, et al. Synthesis without meta-analysis (SWiM) in systematic reviews: reporting guideline. BMJ. 2020;368:l6890.
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA. Chapter 10: Analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Friede T, Rover C, Wandel S, Neuenschwander B. Meta-analysis of few small studies in orphan diseases. Res Synth Methods. 2017;8(1):79-91.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71