Meta-analysis for medical and health sciences

We support evidence synthesis in medicine and the health sciences: meta-analyses and systematic reviews of randomized trials, cohort studies, diagnostic accuracy studies and prognostic research, planned and reported to the standards that clinical journals expect. The work covers protocol and study design, statistical synthesis, risk-of-bias and certainty assessment, and the manuscript.

Evidence synthesis in medicine and health sciences

Medicine is the field in which systematic review and meta-analysis methods were first developed and are most formalized. Questions about treatments, tests, prognosis and risk are usually addressed by many studies of moderate size that do not agree, and decisions about care, policy and further research depend on a defensible summary of that evidence. Organizations such as Cochrane helped establish the methods that most clinical journals now expect, and the guidance in the Cochrane Handbook remains the main reference for reviews of interventions.

Several features distinguish health research from other disciplines. Outcomes are chosen for their importance to patients, so they are often events or times to events and not continuous measurements. Study designs range from randomized trials to registry analyses, each with its own threats to validity. Reporting is standardized through PRISMA 2020 and its extensions, protocols are expected to be registered in advance, and the certainty of the evidence is rated in a structured way using GRADE. The consequences of a wrong conclusion are also greater, which is why the conventions in this field are more prescriptive than in many others.

Medical and health sciences are the flagship area of this service. The same statistical framework described under meta-analysis and systematic review applies, with the design choices, outcome measures and appraisal tools specific to clinical and public health research described on this page.

This service provides research and evidence-synthesis support. It does not provide clinical advice, and the results of a review are not patient-specific guidance.

Specialties and the evidence they rely on

Each specialty has its own typical study designs, outcome measures and recurring methodological problems. The summaries below indicate what distinguishes the evidence in each area, and each links to a page that develops the specialty in detail.

Oncology

Overall and progression-free survival, hazard ratios, and the question of whether surrogate endpoints predict survival.

Cardiology

Composite cardiovascular endpoints, antithrombotic and device trials, and time-to-event outcomes.

Neurology

Stroke and dementia trials, ordinal functional scales such as the modified Rankin Scale, and small-trial heterogeneity.

Psychiatry

Symptom-scale effect sizes, placebo response, drug and psychotherapy comparisons, and dropout as an outcome.

Pediatrics

Neonatal and child outcomes, age-stratified analyses and long-term follow-up.

Surgery

Randomized and observational surgical evidence, learning-curve bias, limited blinding and complication rates.

Obstetrics & gynecology

Maternal and neonatal outcomes, pregnancy-specific endpoints and observational exposures.

Orthopedics

Patient-reported outcome measures, minimal important differences, implants and rehabilitation.

Infectious diseases

Vaccine efficacy, antimicrobial trials, and prevalence and seroprevalence synthesis.

Nursing

Nursing interventions, patient-reported and care-process outcomes, integrative and qualitative synthesis.

Pharmacy

Drug efficacy and safety, rare adverse events and adherence interventions.

Dentistry

Split-mouth and clustered designs, periodontal and implant outcomes, and unit-of-analysis problems.

Nutrition

Dietary exposures and dose-response relations, observational versus trial evidence, and biomarker outcomes.

Public health & epidemiology

Prevalence and incidence synthesis, observational designs and MOOSE reporting.

Rehabilitation & physiotherapy

Functional outcome scales, variation in exercise dose and small trials.

Study designs and what each contributes

A health-related synthesis usually draws on more than one kind of study, and the design determines both the appraisal tool and the way results are pooled.

Randomized trials
The primary evidence for the effects of an intervention. Risk of bias is assessed with RoB 2, and in GRADE the evidence starts at high certainty. Common problems are small trials, differences between comparators, and selective reporting of outcomes.
Cohort and case-control studies
Needed for harms, long-term outcomes and exposures that cannot be randomized. They are pooled as adjusted estimates and appraised with ROBINS-I or the Newcastle-Ottawa Scale, with reporting following MOOSE. Confounding is the central threat.
Diagnostic accuracy studies
Estimate sensitivity and specificity against a reference standard. Appraised with QUADAS-2 and pooled with bivariate models, as described under diagnostic accuracy meta-analysis.
Prognostic and prediction studies
Examine which factors or models predict outcomes. Prognostic factor studies are appraised with QUIPS and prediction model studies with PROBAST. See prognostic meta-analysis.
Prevalence and incidence studies
Estimate how common a condition is. The quality of the sampling frame and large between-study heterogeneity are the main issues. See prevalence meta-analysis.
Qualitative and mixed-methods studies
Describe the experiences and views of patients and clinicians. They are synthesized with methods such as thematic synthesis and reported with guidance such as ENTREQ. See qualitative evidence synthesis.

Common research questions

Most health-related reviews fall into a small number of question types. Each points to a different method.

  • Does an intervention work compared with a control? Pairwise meta-analysis of randomized trials, usually with a random-effects model. See pairwise meta-analysis.
  • Which of several treatments works best? Network meta-analysis, which combines direct and indirect comparisons and requires a plausible transitivity assumption.
  • How accurate is a diagnostic or screening test? Meta-analysis of sensitivity and specificity, with attention to the threshold used and the patient spectrum.
  • Which factors predict an outcome? Synthesis of prognostic factor or prediction model studies, typically using hazard ratios and discrimination measures.
  • How common is a condition, and how does it vary? Meta-analysis of proportions or rates, with exploration of heterogeneity by region, period and population.
  • Is a treatment or exposure safe? Synthesis of adverse events, often rare, from trials and observational studies, using methods suited to sparse data.
  • Does risk change with the amount of an exposure? Dose-response meta-analysis, common in nutrition and environmental epidemiology.
  • How do patients and clinicians experience care? Qualitative or mixed-methods synthesis, often in nursing and public health.

Methods commonly used in health research

The table links the common question types to the method, the usual effect measure and the reporting standard that journals look for.

Question type, method and reporting standard
QuestionMethodTypical effect measureReporting standard
Effect of an interventionPairwise meta-analysisRisk ratio, odds ratio, hazard ratio, mean differencePRISMA 2020
Comparison of many treatmentsNetwork meta-analysisRelative effects and treatment rankingsPRISMA extension for network meta-analysis
Test accuracyDiagnostic accuracy meta-analysisSensitivity and specificityPRISMA-DTA; QUADAS-2 for appraisal
PrognosisPrognostic synthesisHazard ratio, discrimination measuresPRISMA 2020; QUIPS or PROBAST for appraisal
Prevalence or incidenceProportion or rate meta-analysisPooled proportion or ratePRISMA 2020; MOOSE for observational data
Exposure-response relationDose-response meta-analysisRelative risk per unit of exposureMOOSE
Participant-level questionsIndividual participant data meta-analysisParticipant-level treatment effectsPRISMA-IPD

Outcomes and effect measures in health research

Choosing and interpreting the outcome is often harder in health research than the statistics that follow. Five issues come up repeatedly.

Binary outcomes and baseline risk

Risk ratios and odds ratios are relative measures, and a relative effect that looks consistent across trials can correspond to very different absolute effects when the baseline risk differs. A useful report therefore converts the pooled relative effect into absolute terms for stated baseline risks, so that readers can see how many events would be expected to change in a population like their own. For common outcomes the odds ratio overstates the risk ratio, and it should not be described as if it were one.

Time-to-event outcomes

Survival and time-to-recurrence outcomes are pooled as hazard ratios on the log scale. Not every trial reports the hazard ratio and its standard error. When they are missing, they can sometimes be derived from reported confidence intervals, from log-rank statistics, or estimated from published survival curves, and the reliability of each route differs. The proportional hazards assumption should be examined when follow-up is long or curves cross.

Continuous outcomes and patient-reported measures

Symptom and function outcomes are measured on many different scales, which is why standardized mean differences are common in psychiatry, orthopedics and rehabilitation. A standardized effect is hard to interpret clinically, so results are usually re-expressed on a familiar scale and compared with the minimal important difference where one has been established. Mixing change-from-baseline and final-value data without care can distort results.

Composite and surrogate endpoints

Composite endpoints, such as major adverse cardiovascular events, combine outcomes of different importance and frequency, so a favorable composite can be driven by the least important component. The composite and its components should both be examined. Surrogate endpoints, such as progression-free survival or a laboratory marker, are acceptable substitutes only to the extent that they are shown to predict outcomes that matter to patients, and trial-level analyses in oncology have found the strength of that association to vary widely.

Adverse events and rare events

Harms are reported inconsistently and are often too rare for standard inverse-variance methods. Mantel-Haenszel or Peto approaches, careful handling of studies with zero events, and the use of observational data for long-term safety are part of any serious analysis of harms. Absence of reported events is not evidence of safety when follow-up is short or events were not systematically sought.

Methodological considerations specific to health evidence

Risk of bias and certainty of evidence

Appraisal tools are matched to design: RoB 2 for randomized trials, ROBINS-I for non-randomized studies of interventions, QUADAS-2 for diagnostic accuracy, and QUIPS or PROBAST for prognostic work. GRADE then rates the certainty of the body of evidence for each outcome. Five factors can lower certainty (risk of bias, inconsistency, indirectness, imprecision and publication bias), and three can raise it for observational evidence (a large effect, a dose-response gradient, and residual confounding that would reduce the observed effect). The service for risk-of-bias and certainty assessment covers this stage.

Protocol registration and selective reporting

A protocol written before the review begins, and registered where an appropriate registry exists, limits the freedom to change eligibility criteria or analyses once results are seen. PROSPERO accepts systematic reviews with a health-related outcome. Within the included studies, selective outcome reporting is a recognized source of bias, which is why reviewers compare the outcomes reported in a publication with those pre-specified in the trial registry or protocol whenever these are available.

Duplicate publications and overlapping data

One trial can generate several papers, and related reviews can include the same trials. Counting a trial twice inflates its weight and narrows the confidence interval. Linking reports to their parent study, and checking for overlap in umbrella reviews, is routine work in health evidence synthesis.

Populations, baseline risk and clinical heterogeneity

Differences in patient characteristics, severity, setting, co-interventions and follow-up often explain more of the variation between studies than chance does. Subgroup analysis and meta-regression can explore this, but only a limited number of study-level characteristics can be examined, and findings should be treated as hypotheses. Network meta-analysis adds the further requirement that the populations compared indirectly are sufficiently similar for the transitivity assumption to hold.

Unit of analysis

Cluster-randomized trials, crossover trials, split-mouth designs and trials with more than two arms cannot be entered as independent two-arm comparisons without adjustment. Unadjusted entry gives intervals that are too narrow, and the correction depends on information that is often not reported in the paper.

Existing reviews

Many topics already have several systematic reviews, some of them redundant or of low quality. Before a new review is planned, existing reviews are searched for, appraised and used to define what a new one would add. A well-justified update or a review with a new question is more useful than another review of the same trials.

How we support health research projects

Support

From question to submission in health research

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A structured question, eligibility criteria and an analysis plan, with registration prepared where appropriate.

  • Searching and appraisal

    Search strategies for the relevant databases and registries, risk-of-bias assessment by design and a GRADE assessment for each outcome.

  • Synthesis

    Pairwise, network, diagnostic accuracy or prognostic analysis, with sensitivity, subgroup and small-study analyses.

  • Manuscript and submission

    Reporting-guideline checklists, the manuscript and the preparation of submission materials.

Get a quoteDescribe your question, study type and target journal.

Boundaries of this service

Evidence synthesis summarizes what published and registered studies show. It does not replace clinical judgment, and the results of a review do not tell an individual patient what to do. Developing clinical practice guidelines, which involves weighing benefits, harms, values and resources, is a separate process with its own methods and panels, and falls outside the scope described here. We also do not provide advice on the care of individual patients or on the interpretation of an individual's test results.

Where a question is better answered by a different design, such as a scoping review to map the evidence or a narrative synthesis because the studies cannot be pooled, we say so at the outset.

Frequently asked questions

Which reporting guideline should a health-related systematic review follow?

PRISMA 2020 is the standard guideline for systematic reviews, with extensions for diagnostic accuracy reviews, network meta-analyses, scoping reviews and individual participant data. Meta-analyses of observational studies are often also reported against MOOSE. The instructions for authors of the target journal indicate which are required.

Can a meta-analysis include both randomized trials and observational studies?

It can, but they are usually analyzed separately or in clearly labeled subgroups. Observational studies are subject to confounding that randomization avoids, so mixing them in a single pooled estimate can hide differences between the designs. The decision depends on the question and should be stated in the protocol.

Do I need to register my review in PROSPERO?

Registration is strongly encouraged, and many journals ask for it. PROSPERO accepts systematic reviews with a health-related outcome. Reviews that do not meet its criteria, such as scoping reviews, can often be registered on other platforms. Check the current eligibility rules of the registry and the requirements of your target journal.

How is the certainty of evidence assessed?

Most health-related reviews use GRADE, which rates the certainty of the evidence for each outcome as high, moderate, low or very low. Certainty can be lowered for risk of bias, inconsistency, indirectness, imprecision and publication bias, and raised for large effects, dose-response gradients or confounding that would work against the observed effect.

Can you help with network meta-analysis of several treatments?

Yes, network meta-analysis is part of the methods described on this site. It needs a connected network of comparisons and a transitivity assumption that is plausible for the populations involved, and the certainty of the results is assessed with network-specific approaches. A feasibility check at the start establishes whether a network meta-analysis is suitable.

Does this service provide medical or clinical advice?

No. It provides research and evidence-synthesis support. The results of a review or analysis are not advice about the care of an individual patient.

References

  1. Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
  2. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
  3. Hutton B, Salanti G, Caldwell DM, et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions. Ann Intern Med. 2015;162(11):777-784.
  4. McInnes MDF, Moher D, Thombs BD, et al. Preferred reporting items for a systematic review and meta-analysis of diagnostic test accuracy studies: the PRISMA-DTA statement. JAMA. 2018;319(4):388-396.
  5. Stewart LA, Clarke M, Rovers M, et al. Preferred reporting items for systematic review and meta-analyses of individual participant data: the PRISMA-IPD statement. JAMA. 2015;313(16):1657-1665.
  6. Stroup DF, Berlin JA, Morton SC, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-2012.
  7. Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924-926.
  8. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  9. Sterne JA, Hernan MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
  10. Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
  11. Hayden JA, van der Windt DA, Cartwright JL, Cote P, Bombardier C. Assessing bias in studies of prognostic factors. Ann Intern Med. 2013;158(4):280-286.
  12. Wolff RF, Moons KGM, Riley RD, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58.
  13. Tong A, Flemming K, McInnes E, Oliver S, Craig J. Enhancing transparency in reporting the synthesis of qualitative research: ENTREQ. BMC Med Res Methodol. 2012;12:181.
  14. Reitsma JB, Glas AS, Rutjes AWS, Scholten RJPM, Bossuyt PM, Zwinderman AH. Bivariate analysis of sensitivity and specificity produces informative summary measures in diagnostic reviews. J Clin Epidemiol. 2005;58(10):982-990.
  15. Parmar MKB, Torri V, Stewart L. Extracting summary statistics to perform meta-analyses of the published literature for survival endpoints. Stat Med. 1998;17(24):2815-2834.
  16. Tierney JF, Stewart LA, Ghersi D, Burdett S, Sydes MR. Practical methods for incorporating summary time-to-event data into meta-analysis. Trials. 2007;8:16.
  17. Prasad V, Kim C, Burotto M, Vandross A. The strength of association between surrogate end points and survival in oncology: a systematic review of trial-level meta-analyses. JAMA Intern Med. 2015;175(8):1389-1398.
  18. Page MJ, McKenzie JE, Kirkham J, et al. Bias due to selective inclusion and reporting of outcomes and analyses in systematic reviews of randomised trials of healthcare interventions. Cochrane Database Syst Rev. 2014;(10):MR000035.
  19. Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082.
  20. Ioannidis JPA. The mass production of redundant, misleading, and conflicted systematic reviews and meta-analyses. Milbank Q. 2016;94(3):485-514.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.