What a meta-analysis is, and what it is not
A meta-analysis is a statistical method for combining the quantitative results of two or more separate studies that address a sufficiently similar question. Rather than reading each study on its own, it estimates an overall effect, measures how much the studies disagree with one another, and examines whether that disagreement can be explained by differences in populations, interventions, exposures, outcomes or study quality.
When the studies were identified through a systematic, documented search, the pooled estimate describes the body of available evidence rather than only the studies that were easiest to find. That is why a meta-analysis is normally carried out inside a systematic review. The two terms are not interchangeable. The review is the structured process of identifying, appraising and summarizing studies; the meta-analysis is an optional statistical step within it. A review can be sound without a meta-analysis when the studies are too different to combine, and a meta-analysis is only as reliable as the search and selection behind it. The guide on meta-analysis versus systematic review sets out the distinction in more detail.
A meta-analysis is also not a way of rescuing weak evidence. Combining studies that share the same bias produces a precise but biased answer. Appraisal of risk of bias and a judgment about the certainty of the evidence therefore belong alongside the statistics, not after them.
When a meta-analysis is appropriate
Whether a meta-analysis is appropriate depends on the research question, the studies that exist and the data they report, not on the discipline alone. Four conditions usually need to hold.
- The studies ask comparable questions. Population, intervention or exposure, comparator and outcome must be similar enough that an average is meaningful rather than an arithmetic accident.
- The outcomes can be placed on a common scale. Either the studies measure the outcome in the same way, or the results can be converted to a shared effect measure such as a standardized mean difference.
- The studies were found systematically. A documented search and explicit eligibility criteria guard against pooling only what was convenient or favorable.
- Usable data are available. Each study must report an effect estimate with a measure of its precision, or the raw data needed to compute one.
There is no fixed minimum number of studies. Two studies can be combined mathematically, but with few studies the estimate of between-study variation is imprecise and several checks, including tests for small-study effects, are unreliable. The table below summarizes how common situations are usually handled.
| Situation | Usual approach | Main consideration |
|---|---|---|
| Several randomized trials with the same outcome | Pairwise meta-analysis, usually with a random-effects model | Risk of bias in the trials and consistency of the comparator |
| Observational studies reporting adjusted estimates | Generic inverse-variance pooling of adjusted estimates, reported against MOOSE | Confounding and different adjustment sets across studies |
| Many interventions connected through shared comparators | Network meta-analysis | Whether the transitivity assumption is plausible |
| Studies of diagnostic test accuracy | Diagnostic accuracy meta-analysis with bivariate models | Threshold effects and spectrum of patients |
| Studies that differ widely in design, population or outcome | Structured synthesis without pooling, or a scoping review | A single average would hide more than it shows |
| One study, or two studies that clearly disagree | Report the studies individually and explain why they differ | Pooling adds little and can mislead |
What the service includes
The core of the service is the statistical analysis, from agreeing the analysis plan to writing the results. It can follow a review you have already conducted, or form one stage of a larger project. When agreed at the scoping stage, it extends to preparing the complete manuscript and the materials for journal submission.
- Feasibility check. A review of the question, the included studies and the available data to confirm that pooling is justified and which approach fits.
- Analysis plan. The effect measure, model, heterogeneity assessment, subgroup and sensitivity analyses specified before the results are examined.
- Data audit and effect-size computation. Checking the extraction dataset for errors, converting results to a common scale, and handling multi-arm, cluster and repeated-measures designs correctly.
- Statistical analysis. Fixed-effect and random-effects models as justified, heterogeneity statistics, prediction intervals, subgroup analysis and meta-regression where the data allow.
- Small-study effects and publication bias. Funnel plots and appropriate tests when enough studies are available, with sensitivity analyses to show how much conclusions depend on them.
- Certainty of evidence. Support with a GRADE assessment where it is within the scope of the project.
- Figures and tables. Forest plots, funnel plots and summary tables formatted to journal standards.
- Methods and results text. Written sections aligned with PRISMA 2020 or MOOSE, including a description of the software and settings used.
- Revision support. Responses to statistical queries from editors and reviewers, including re-analysis when requested.
- Full manuscript preparation (optional). Drafting or revising the complete paper, including the abstract, introduction, methods, results, discussion, tables and reference list, built on the analysis and on the authors' own interpretation of the findings.
- Submission support (optional). Journal selection, formatting to the instructions for authors, the completed reporting-guideline checklist and flow diagram, a cover letter, and preparation of the submission files. See publication support and manuscript editing.
Literature searching, screening and risk-of-bias appraisal are separate stages of a review. They are described under systematic review, search strategy development and risk-of-bias and certainty assessment.
What data a meta-analysis requires
The data requirements follow from the outcome type. Each study contributes a result and a measure of how precisely that result was estimated. Where a study does not report them directly, they can often be derived from other reported quantities, with the conversion documented.
- Binary outcomes
- Events and total participants in each group, supporting risk ratio, odds ratio or risk difference.
- Continuous outcomes
- Mean, standard deviation and number of participants per group, supporting mean difference or standardized mean difference. Change-from-baseline and final values should not be mixed without care.
- Time-to-event outcomes
- Log hazard ratios and their standard errors, taken from the publication or derived from reported confidence intervals. Reconstruction from survival curves is possible but less reliable.
- Correlations
- The correlation coefficient and sample size, usually pooled after Fisher's z transformation.
- Proportions
- Number of events and sample size, used for prevalence and incidence, usually pooled after a variance-stabilizing or logit transformation.
- Adjusted estimates
- An effect estimate and standard error from each study, pooled with the generic inverse-variance method. This is the usual route for observational studies.
Beyond the outcome data, an analysis that explores differences between studies needs study-level characteristics such as population, setting, dose or follow-up, and, where an appraisal has been done, the risk-of-bias judgments. The unit of analysis must also be clear. A trial with three arms, a cluster-randomized trial or a study reporting several outcomes at several time points cannot simply be entered as independent rows.
When data are missing, the first step is to look for them in the full text, supplements and registry records, and to contact the study authors where this is appropriate. Imputing a standard deviation or converting a median and range to a mean is sometimes acceptable, but each such step should be reported and tested in a sensitivity analysis.
How the analysis is performed
A meta-analysis is a sequence of decisions, and each decision should be made for a stated reason and, ideally, before the results are seen. The main ones are described below. The statistical choices follow the guidance of the Cochrane Handbook and the wider methodological literature listed under References.
Choosing the effect measure
The effect measure determines what the pooled number means. Risk ratios are easier to interpret than odds ratios for common outcomes, while odds ratios have convenient statistical properties and are the natural output of case-control studies and logistic models. Standardized mean differences allow studies that used different scales for the same construct to be combined, at the cost of results that depend on the variability of each sample. The measure is chosen for interpretability and for the data available, and alternatives are examined as a sensitivity analysis. See effect sizes in meta-analysis.
Choosing and justifying the model
A fixed-effect model assumes one true effect shared by all studies, so differences between them are attributed to chance. A random-effects model assumes that true effects vary and estimates their mean and spread. In most applied settings the studies differ in ways that make the second assumption more credible, but the choice should be reasoned, not automatic. The page on fixed-effect and random-effects models explains the difference, and the method page on random-effects meta-analysis covers estimation.
Within the random-effects framework, the estimator of between-study variance matters. The DerSimonian-Laird estimator has long been the default in software, but comparisons in the methodological literature favor alternatives such as restricted maximum likelihood or Paule-Mandel in many settings. With few studies, the Hartung-Knapp-Sidik-Jonkman adjustment to the confidence interval generally gives more realistic coverage than the conventional Wald interval. For sparse binary data, with few events or zero cells, Mantel-Haenszel or Peto methods and carefully chosen continuity corrections avoid problems that inverse-variance weighting can create. Where the choice affects the conclusion, results from more than one approach are shown.
Quantifying heterogeneity
Heterogeneity is variation in the true effects across studies beyond what chance would produce. It is summarized with the Q statistic, the between-study variance (τ²) and the I² statistic, which expresses the share of observed variability attributable to heterogeneity and not to sampling error. I² depends on the precision of the included studies, so it should not be read as an absolute measure of how different the studies are, and the Q test has low power when studies are few. A prediction interval is often more informative than any of these, because it describes the range in which the true effect of a new study would be expected to fall. Background is given in the guides on heterogeneity and I².
Exploring why studies differ
When heterogeneity is present, the next question is whether study characteristics explain it. Subgroup analysis compares pooled estimates across categories and should rely on a formal test for interaction, not on whether each subgroup is individually significant. Meta-regression relates effect sizes to continuous or categorical moderators. Both are observational comparisons between studies, so they can suggest explanations but cannot prove them, and a commonly cited guide is to have at least ten studies for each characteristic modeled. Associations at the study level also do not necessarily apply to individual participants.
Sensitivity analyses
Sensitivity analyses test whether the conclusion survives reasonable changes in method. Typical examples are leaving out one study at a time, excluding studies at high risk of bias, using a different estimator or effect measure, and removing studies whose data required substantial conversion. Influence diagnostics identify studies that move the pooled estimate disproportionately. The aim is not to find the most favorable result but to report how robust the finding is. See sensitivity analysis in meta-analysis.
Small-study effects and publication bias
Studies with small samples or statistically significant results are more likely to be published and found. A funnel plot, which graphs each study's effect against its precision, can reveal asymmetry that may indicate this problem, although asymmetry can also arise from real differences between small and large studies. Formal tests such as Egger's regression are recommended only when about ten or more studies are available, and for some effect measures, notably the log odds ratio, alternative tests perform better. Methods such as trim-and-fill and selection models are used as sensitivity analyses, not as corrections to be trusted on their own. See funnel plots and publication bias.
Dependent effect sizes
Standard meta-analytic models assume that each effect size is independent. That assumption fails when one study contributes several outcomes, time points or comparisons, or when studies share a research team, sample or control group. Depending on the structure, the appropriate remedy may be to select one estimate by a pre-specified rule, to use a multilevel model, to apply robust variance estimation, or to fit a multivariate meta-analysis. Ignoring dependence gives confidence intervals that are too narrow.
Study designs and outcome types
The same underlying framework applies across several study designs, with adjustments for each.
- Randomized trials are pooled using the risk of bias assessments from tools such as RoB 2, and the certainty of evidence is rated with GRADE.
- Observational studies are usually pooled as adjusted estimates. Confounding is the central threat, so the extent and quality of adjustment are examined, and reporting follows MOOSE.
- Prevalence and incidence studies pool proportions or rates, where between-study heterogeneity is usually large and the quality of the sampling frame matters. See prevalence meta-analysis.
- Correlational and survey-based studies, common in psychology and management, pool correlation coefficients and examine moderators, often with corrections for measurement reliability.
- Experimental studies in ecology and agriculture commonly use ratio-based effect sizes and must account for non-independence among comparisons from the same site or species.
More specialized designs have their own pages: network meta-analysis, Bayesian meta-analysis, dose-response meta-analysis, prognostic meta-analysis and individual participant data meta-analysis. Not every method suits every field or every dataset, and part of the scoping work is to say so.
How a project runs
Scope and feasibility
You describe the question, the included studies and the target journal. We assess whether pooling is justified, what data exist, and which approach fits, and we agree what is in and out of scope.
Analysis plan
The effect measure, model, handling of dependent data, and planned subgroup and sensitivity analyses are written down before the analysis begins, so that later choices can be distinguished from planned ones.
Data audit and coding
The extraction dataset is checked against the source papers for a sample or in full, depending on the agreement. Effect sizes are computed, conversions are documented, and unusual values are queried.
Analysis and review
The planned analyses are run, results are reviewed for plausibility and consistency, and any deviation from the plan is recorded with its reason.
Reporting and revision
Figures, tables and the methods and results text are prepared against the relevant reporting guideline. Where agreed, the full manuscript and submission materials follow, with support through journal review and revision.
Get a quoteShare your scope and deadline, and we will propose what is feasible.
Deliverables
The package depends on the agreed scope. A full meta-analysis project typically includes the following.
- Analysis plan describing every planned analysis and its justification.
- Cleaned analysis dataset with effect sizes, standard errors and the conversions applied.
- Statistical report with pooled estimates, confidence and prediction intervals, heterogeneity statistics, and sensitivity and subgroup results.
- Figures in publication-ready formats, including forest plots and, where appropriate, funnel plots.
- Methods and results text written for the manuscript and aligned with the reporting guideline.
- Analysis code and output logs so that the results can be reproduced by someone else.
- Response support for statistical comments from reviewers.
Full manuscript and submission package
When the scope includes manuscript preparation and submission, the package also contains the following.
Full manuscript draft
Abstract, introduction, methods, results, discussion, tables and formatted references.
Reporting-guideline checklist
The completed checklist and a PRISMA flow diagram, checked against the final text.
Submission package
A journal recommendation, the manuscript formatted to that journal, a cover letter and supplementary files.
Revision round
A point-by-point response to editor and reviewer comments, with re-analysis where requested.
Manuscript preparation follows the research integrity and authorship statement. The researchers who conceived the study and interpret its findings remain responsible for the content and its conclusions, and contributions that do not meet authorship criteria are acknowledged.
Quality assurance and reproducibility
A meta-analysis should be checkable. The working practices below are designed to make that possible.
- Pre-specification. The analysis plan is fixed before results are examined, and deviations are listed with reasons.
- Verification of inputs. Extracted values are checked against the source documents, because transcription and unit errors are among the most common causes of wrong pooled estimates.
- Independent re-analysis. Key results are re-computed with a second method or tool to detect coding errors.
- Documented software. The software, package versions and settings are recorded and delivered with the code.
- Confidential handling. Unpublished manuscripts and data are treated as confidential, as set out in the research integrity statement.
Limitations, and when we advise against pooling
A meta-analysis cannot be better than the studies it combines. Biased or poorly reported studies yield a biased pooled estimate, however sophisticated the model. Several other limitations deserve plain statements in any report.
- Substantial heterogeneity may mean that a single pooled number is not a useful summary, even when a random-effects model is used.
- With few studies, estimates of between-study variance are imprecise, confidence intervals can be too narrow, and tests for small-study effects have little power.
- Pooled observational estimates remain vulnerable to confounding, and a large number of studies does not remove it.
- Subgroup and meta-regression results are associations between studies and should be treated as hypotheses.
- Statistical significance of a pooled result is not the same as importance. The size of the effect and its interval need to be interpreted against the context.
We advise against pooling when the studies ask clearly different questions, when outcome definitions cannot be reconciled, when most of the data would have to be estimated, or when the studies are at such high risk of bias that a pooled number would give false reassurance. In those cases a structured narrative synthesis, reported according to the SWiM guideline, or a different review design is more honest. See methods for alternatives.
This service provides research and evidence-synthesis support. It does not provide clinical advice, and results are not patient-specific guidance.
How long does a meta-analysis take?
The time required depends on the number of studies, how much of the data must be derived or converted, the number of outcomes and subgroups, whether searching and screening are part of the project, and how many rounds of revision follow. A review with a few homogeneous trials and one outcome is a different task from one with dozens of studies, several outcomes and dependent effect sizes. After the feasibility check, the timeline and the steps it depends on are discussed and agreed so that the deadline is realistic before work begins. If a journal deadline is fixed, tell us at the start so that the scope can be set accordingly.
Meta-analysis across disciplines
The statistical framework is shared, but the effect measures, study designs and conventions differ by field.
- Medical and health sciences work mostly with randomized trials and cohort studies, binary and time-to-event outcomes, and GRADE for certainty.
- Psychology and behavioral sciences often pool standardized mean differences and correlations, and must consider measurement reliability and replicability.
- Management and business research frequently combines correlations across organizations and examines context moderators.
- Economics uses meta-regression analysis on estimates such as elasticities, with explicit attention to publication selection and specification differences.
- Education research deals with clustered designs and learning outcomes measured on varied tests.
- Life sciences and ecology pool experimental and observational effect sizes where non-independence is a standing concern.
Frequently asked questions
How many studies do I need for a meta-analysis?
There is no universal minimum. Two studies can be combined mathematically, but with few studies the estimate of between-study variation is imprecise and tests for small-study effects are unreliable. Whether pooling is meaningful depends on how similar the studies are as much as on how many there are.
Should I use a fixed-effect or a random-effects model?
A fixed-effect model assumes one true effect shared by all studies, and a random-effects model allows the true effect to vary. In most applied research the studies differ enough that random effects is the more credible assumption, but the choice should be justified and checked in a sensitivity analysis.
Can you perform a meta-analysis of observational studies?
Yes, with caution. Observational studies are usually pooled as adjusted estimates and reported against the MOOSE guideline. Confounding and differences in adjustment between studies are the main concerns, and they limit how the result can be interpreted.
What if my meta-analysis is not statistically significant, or the studies are very heterogeneous?
A non-significant result is still a result and should be reported with its confidence interval. High heterogeneity is explored through subgroup analysis or meta-regression where the data allow, and the report states plainly when a single pooled estimate is not a useful summary.
Do you also run the literature search and screening?
Searching, screening and risk-of-bias assessment are separate stages of a systematic review and are described on their own service pages. A meta-analysis can begin from an extraction dataset you already have, after a data audit.
Which software is used for the analysis?
The software is agreed at the scoping stage. R, Stata and RevMan are all widely used and accepted by journals, and the software and versions used are named in the methods text.
Can you prepare the full manuscript and support journal submission?
Yes, as an extension of the analysis when it is agreed at the scoping stage. That can include the complete draft, the reporting checklist, journal selection, formatting, a cover letter and support through revision. The authors remain responsible for the content, interpretation and conclusions, and authorship follows the research integrity statement.
Is this service clinical advice?
No. It provides research and evidence-synthesis support. The results of a meta-analysis are not patient-specific medical advice.
References
- Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
- Stroup DF, Berlin JA, Morton SC, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-2012.
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
- Cooper H, Hedges LV, Valentine JC, editors. The Handbook of Research Synthesis and Meta-Analysis. 3rd ed. Russell Sage Foundation; 2019.
- DerSimonian R, Laird N. Meta-analysis in clinical trials. Control Clin Trials. 1986;7(3):177-188.
- Higgins JPT, Thompson SG. Quantifying heterogeneity in a meta-analysis. Stat Med. 2002;21(11):1539-1558.
- Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.
- IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6(7):e010247.
- IntHout J, Ioannidis JPA, Borm GF. The Hartung-Knapp-Sidik-Jonkman method for random effects meta-analysis is straightforward and considerably outperforms the standard DerSimonian-Laird method. BMC Med Res Methodol. 2014;14:25.
- Veroniki AA, Jackson D, Viechtbauer W, et al. Methods to estimate the between-study variance and its uncertainty in meta-analysis. Res Synth Methods. 2016;7(1):55-79.
- Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315(7109):629-634.
- Sterne JAC, Sutton AJ, Ioannidis JPA, et al. Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ. 2011;343:d4002.
- Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
- Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924-926.
- Campbell M, McKenzie JE, Sowden A, et al. Synthesis without meta-analysis (SWiM) in systematic reviews: reporting guideline. BMJ. 2020;368:l6890.
- Viechtbauer W. Conducting meta-analyses in R with the metafor package. J Stat Softw. 2010;36(3):1-48.