Guide

Sensitivity analysis in meta-analysis

A sensitivity analysis repeats a meta-analysis with different choices or data to see whether the conclusion survives. It does not search for the most favorable result: its purpose is to show how robust the finding is. This guide describes the common analyses, with a worked leave-one-out example.

What sensitivity analysis is for

Every meta-analysis rests on choices. Which studies to include, which effect measure and model to use, how to handle studies with missing data, which estimator of between-study variance, how to treat dependent estimates: each decision is defensible, and each can influence the result. A sensitivity analysis asks whether the conclusion would be different under reasonable alternatives. If it would not, the finding is robust, and readers can place more trust in it. If it would, the report must say that the conclusion depends on the choice, which is among the most important things a reader can be told.

The purpose should be kept clear. A sensitivity analysis is a test of robustness, not a way of finding a more favorable estimate. The analyses should be specified, as far as possible, in advance, and all of them reported, including those that change the conclusion. A report that shows only the analyses that agree with the main result is misleading, and it defeats the point. The term is used differently in diagnostic testing, where sensitivity is a measure of accuracy, and the two should not be confused. The main method is described on the page for meta-analysis.

Common sensitivity analyses

Leave-one-out
The analysis is repeated omitting each study in turn. It reveals studies that are influential and shows whether the pooled result depends on one of them.
Restricting by risk of bias
Studies at high or unclear risk of bias are excluded, to see whether the conclusion holds in the more trustworthy studies.
Alternative models and estimators
Fixed-effect against random-effects, different estimators of tau squared, and the adjusted against the conventional confidence interval, to see how much the model drives the result.
Alternative effect measures
Odds ratio against risk ratio, or mean difference against standardized difference, where either is defensible.
Excluding studies with derived data
Studies whose effect sizes required substantial conversion or imputation are removed, to test the effect of those assumptions.
Different handling of missing or zero data
Alternative continuity corrections, or a model that avoids them, in sparse-data settings.
Different assumptions about missing studies
Trim-and-fill or selection models, to test sensitivity to possible publication bias.
Different correlation or variance assumptions
Where a correlation or a variance had to be assumed, such as for change scores, a range of plausible values.

A worked leave-one-out analysis

The six simulated studies used elsewhere on this site show the idea. They are invented for illustration. With all six studies, the random-effects pooled estimate is 0.22, with a 95 percent confidence interval from 0.10 to 0.35. The table shows the result when each study is omitted in turn.

Leave-one-out results (simulated data, random effects)
Omitted studyPooled estimate95% CI
None (all six)0.220.10 to 0.35
Study A0.210.06 to 0.36
Study B0.260.12 to 0.40
Study C0.200.08 to 0.33
Study D0.230.06 to 0.40
Study E0.250.14 to 0.37
Study F0.190.06 to 0.32

The pooled estimate ranges from 0.19 to 0.26 across the six omissions, so no single study changes it much, and the interval stays above zero each time or nearly so, which points to a stable conclusion. A second analysis restricts to the four studies assessed at low risk of bias, giving 0.17 with an interval from 0.04 to 0.29. The estimate is lower than in the full set, and the interval is wider because there are fewer studies, which is the usual trade-off: restriction removes the less trustworthy data and costs precision. Reporting both lets the reader see that the estimate is similar in the better studies, and that the conclusion depends on how much precision one is willing to give up. The risk-of-bias labels are invented.

Alternative models and estimators

The commonest sensitivity analysis of the model asks whether the result depends on the choice between fixed-effect and random-effects analysis, and on the estimator and interval method within random effects. The analyst re-runs the analysis with the alternatives and compares the estimate, the interval and the heterogeneity statistics. If the estimator of tau squared matters, for example because the DerSimonian-Laird estimate is much smaller than the restricted maximum likelihood estimate, the conclusion may differ in the width of the interval, and the report should explain which is preferred and why. With few studies the adjusted Hartung-Knapp-Sidik-Jonkman interval is compared with the conventional one. For sparse binary data, a Mantel-Haenszel or exact analysis is compared with an inverse-variance analysis with a continuity correction. Bayesian analyses add a sensitivity analysis to the prior. Agreement across these choices is reassuring, and disagreement tells the reader that the model matters. See the page on random-effects meta-analysis.

Dependent data and imputed values

Some choices are forced by incomplete reporting. When a study contributes several correlated effect sizes, the analyst must choose one by a rule or model the dependence, and the chosen rule, for example taking the effect at the longest follow-up, can be varied. When a standard deviation had to be imputed from other studies, a correlation had to be assumed for change scores, or a median and range had to be converted to a mean and standard deviation, the sensitivity analysis varies the assumed value over a plausible range, or excludes the studies concerned, to see how much the result moves. These analyses are often the most informative of all, because they test assumptions that cannot be checked from the data, and they provide readers with the information needed to judge whether the assumption is important. A conclusion that changes when a plausible correlation changes from 0.3 to 0.7 should be reported as sensitive to that assumption.

What sensitivity analysis cannot do

Sensitivity analysis shows the effect of the choices made, not the effect of choices not considered. It cannot reveal bias shared by all the studies, and it cannot correct for it. A conclusion that is robust to every analysis tried may still be wrong if the underlying studies are biased in the same direction, or if relevant studies are missing. It also cannot rescue an analysis with too little data: with a handful of studies, every analysis is imprecise, and apparent stability may reflect a lack of information. It should be interpreted together with the assessment of risk of bias, heterogeneity and reporting bias, and with the certainty of evidence, and not in place of them.

Influence diagnostics

Leave-one-out analysis is the simplest influence diagnostic. More formal measures quantify how much each study changes the pooled estimate, the heterogeneity and the model fit. Standardized residuals show which studies are outliers relative to the model, and measures such as Cook's distance, the change in the estimate scaled by its precision, and the covariance ratio identify influential points. Graphical tools, such as a Baujat plot, which graphs each study's contribution to heterogeneity against its influence on the result, help to locate studies that matter. An influential study is not necessarily an error. It may be large and precise, or different for a real reason. The step after finding one is to investigate it: check the extracted data, examine its design and population, and consider whether it should be treated separately. Dropping it because it is inconvenient is not acceptable.

Risk of bias and quality

Restricting the analysis to studies at low risk of bias is the sensitivity analysis most often requested by reviewers. It tests whether the conclusion depends on weaker studies. It can be done by excluding studies at high risk overall, or at high risk for a domain such as allocation concealment or blinding, or by running a subgroup analysis by risk of bias. If the effect is smaller in the better studies, that is evidence of bias in the weaker ones and the main conclusion should be framed with care. If it is similar, the finding is reassuring. When most studies are at high risk, the restricted analysis may have too few studies to say much, and this itself is a finding, because it shows how little of the evidence is trustworthy. The assessment methods are described in the guide to risk-of-bias tools.

Planning and prespecification

The most credible sensitivity analyses are specified in the protocol, with a rationale. A short list of analyses chosen because they address the likely threats to validity is better than an open-ended exploration. Typical prespecified analyses include a different model, exclusion of studies at high risk of bias, exclusion of studies with imputed data and an assessment of reporting bias. Further analyses suggested by the data, such as removing an unexpected outlier, are allowed but labeled as post hoc, with the reason. The number of analyses that were performed, and all of their results, are reported. This discipline protects the researchers from the charge of having chosen an analysis to reach a conclusion, and protects readers from a misleading picture.

Reading the results

The results of sensitivity analyses are read as a set. If all agree on the direction and rough size of the effect, the main finding is robust. If one or two change the interval or the significance but not the direction, the finding is moderately robust, and the report notes which choices mattered. If an analysis reverses the conclusion, or makes the effect disappear, the conclusion is fragile, and the report should say so plainly and word its summary accordingly. A fragile result may still be valuable, since it identifies the question to resolve, for example whether a particular study is correct, and it directs future research. In the assessment of certainty, a conclusion that depends on analytic choices contributes to concerns about inconsistency, risk of bias or imprecision. See GRADE.

Reporting

The methods section lists the planned sensitivity analyses and says which were post hoc. The results present the main analysis and each sensitivity analysis in a table or a supplementary figure, with estimates and intervals, and the text summarizes whether the conclusion held. Leave-one-out results are often shown in a plot. The discussion considers the implications of any analysis that changed the conclusion. PRISMA 2020 asks authors to report the methods for sensitivity analyses and the results. See PRISMA 2020.

Common mistakes

  • Running many analyses and reporting only those that support the main result.
  • Dropping an influential study without investigating or reporting it.
  • Treating a sensitivity analysis as a substitute for subgroup analysis, or the reverse.
  • Failing to separate planned and post hoc analyses.
  • Ignoring the loss of precision when restricting to fewer studies.
  • Presenting a fragile conclusion as firm because the main analysis was significant.

Support

Planning, running and reporting sensitivity analyses is part of the meta-analysis service, and responses to reviewers' requests for further analyses are covered by peer-review revision support.

Get a quoteSend your data and any reviewer comments.

Frequently asked questions

What is the difference between sensitivity analysis and subgroup analysis?

A sensitivity analysis tests whether the overall conclusion changes when choices or data are altered. A subgroup analysis compares effects between categories of a characteristic, to explore why studies differ.

How many sensitivity analyses should I do?

A few, chosen in advance because they address the main threats to validity, and all reported. Many analyses chosen after seeing the data invite selective reporting.

Should I remove an outlier?

Investigate it first, by checking the data and the study. If it is an error, correct it. If it differs for a real reason, report the analysis with and without it, and do not simply drop it.

What if the conclusion changes in a sensitivity analysis?

Report that the finding is fragile, say which choice matters, and word the conclusion accordingly.

Is a leave-one-out analysis enough?

It is a useful start, but it does not test choices of model, measure or data handling, which need their own analyses.

Do I need to prespecify sensitivity analyses?

It is strongly recommended. Those not prespecified should be labeled post hoc.

References

  1. Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  2. Viechtbauer W, Cheung MWL. Outlier and influence diagnostics for meta-analysis. Res Synth Methods. 2010;1(2):112-125.
  3. Baujat B, Mahe C, Pignon JP, Hill C. A graphical method for exploring heterogeneity in meta-analyses: application to a meta-analysis of 65 trials. Stat Med. 2002;21(18):2641-2652.
  4. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
  5. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  6. Higgins JPT, Thompson SG, Spiegelhalter DJ. A re-evaluation of random-effects meta-analysis. J R Stat Soc Ser A Stat Soc. 2009;172(1):137-159.
  7. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.