Meta-analysis in oncology

Oncology evidence is dominated by time-to-event outcomes, by trials that report hazard ratios, and by debate over whether early endpoints predict survival. These features shape every step of a review: the effect measure, the data extraction, the handling of crossover and the judgment of certainty.

Evidence synthesis in oncology

Cancer treatment has been shaped by meta-analysis for decades. The Early Breast Cancer Trialists' Collaborative Group pooled individual patient data from randomized trials of adjuvant therapy and showed, by the weight of many trials, effects on recurrence and mortality that no single trial could establish. Today reviews appear for every tumor type and treatment class: chemotherapy regimens, targeted agents, immune checkpoint inhibitors, radiotherapy schedules, surgical approaches and supportive care.

The field has special features. Outcomes are mostly the time to an event, so results arrive as hazard ratios with censored follow-up. Many trials are open-label, which allows assessment bias for outcomes such as progression. Patients often cross over to the experimental treatment after progression, which dilutes the survival difference. New agents are approved on the basis of surrogate endpoints, so there is a continuing question of whether those endpoints predict survival or quality of life. And treatments are given in sequences, so the effect of one drug depends on what comes before and after it.

Methods for these questions are described under meta-analysis and systematic review. The general framework is the one for health sciences, applied here to cancer.

Outcomes and effect measures

Common outcomes in oncology reviews
OutcomeUsual effect measurePoints to check
Overall survivalHazard ratio (log scale pooling)Crossover and subsequent therapy; follow-up length; proportional hazards
Progression-free survivalHazard ratioAssessment schedule, blinded independent review, handling of death and censoring
Objective responseRisk ratio or odds ratioResponse criteria version (such as RECIST); single-arm trials cannot be compared directly
Toxicity (grade 3 or higher)Risk ratio, rate ratioRare events, definitions and reporting thresholds; selective reporting
Quality of lifeStandardized mean difference or change in scale pointsInstrument, missing data from death and dropout, minimal important difference

Reviewers rarely have the log hazard ratio and its standard error given directly. They are derived from the reported hazard ratio and its confidence interval. For example, a hazard ratio of 0.75 with a 95 percent confidence interval of 0.60 to 0.94 has a log hazard ratio of ln(0.75) = -0.288 and a standard error of (ln(0.94) - ln(0.60))/(2 x 1.96) = 0.115. These two values are what the pooling model uses. When only a Kaplan-Meier curve or the number of events is given, methods by Parmar and colleagues and by Tierney and colleagues allow approximate estimation, and the reliability of such estimates should be stated. The guide on hazard ratio meta-analysis explains the steps.

When the hazard ratio is not enough

The hazard ratio assumes the ratio of hazards is constant over time. Immunotherapy trials often show curves that cross or separate late, because the benefit takes time to appear or applies to a subset of patients. A single hazard ratio then averages two different regimes and may misdescribe the result. Reviews can check the assumption through the shape of curves, tests on Schoenfeld residuals where data are available, and by comparing short and long follow-up. Alternatives include the restricted mean survival time, which gives the difference in average survival up to a stated time and has a direct meaning in months, and landmark analyses.

The time frame should be chosen before looking at the data. Reviews that report the difference at several times, with how many patients remain at risk at each, give readers more than a pooled hazard ratio alone. Where individual patient data are available, individual participant data meta-analysis can reconstruct the curves and test the proportional hazards assumption directly.

Surrogate endpoints

Progression-free survival, response rate and disease-free survival are used as endpoints because they can be observed sooner than survival. Whether they are valid substitutes is tested at two levels. At the patient level the question is whether a patient's progression-free survival predicts that patient's survival. At the trial level the question is whether the treatment effect on the surrogate predicts the treatment effect on survival. Trial-level validation uses meta-analysis of trials, typically a regression of the log hazard ratio for overall survival on the log hazard ratio for the surrogate, with the coefficient of determination as the summary of association.

Prasad and colleagues reviewed trial-level analyses in oncology and found that the strength of association between surrogates and survival was often low to moderate, with considerable variation by cancer type and treatment class. A surrogate that works for one class of drugs may not work for another. Good practice is to state the setting in which the surrogate was assessed, to give the uncertainty of the predicted survival effect and not only the correlation, and to avoid using a surrogate validated for one setting in another.

Randomized, single-arm and observational evidence

Randomized trials are the main source for comparative effects. Single-arm trials are common in early-phase studies and for rare cancers, but they have no control group, so their response rates cannot be compared with those of other trials without strong assumptions. Pooling single-arm results gives a summary of response in treated patients, not a treatment effect, and the review should say so. Observational studies and registries describe real-world outcomes in patients who are older and have more comorbidity than trial participants, and they are open to confounding and to immortal time bias.

Network meta-analysis is common in oncology because many regimens compete and few have been tested head to head. It is only as valid as the similarity of the trials. Differences in prior treatment, biomarker status, performance status and the line of therapy can differ among trials and distort indirect comparisons. A review should tabulate these effect modifiers. See network meta-analysis for the assumptions.

Biomarkers, subgroups and prognostic research

Modern oncology depends on biomarkers that predict who benefits from a drug: receptor status, mutations, expression levels. Questions fall into two types. A predictive biomarker modifies the treatment effect, and its evidence needs a test for interaction between marker and treatment within trials, not a comparison of results across marker-positive and marker-negative subgroups from different trials. A prognostic marker is associated with outcome regardless of treatment and is studied through prognostic meta-analysis, where QUIPS and PROBAST help to assess bias.

Subgroup findings in single trials are often unreliable. Reviews should rely on planned subgroup analyses, check interaction tests and, where possible, use data from several trials. Diagnostic questions, such as imaging or liquid biopsy accuracy, are handled with diagnostic accuracy meta-analysis.

Radiotherapy, surgery and supportive care

Not all oncology evidence concerns drugs. Radiotherapy trials compare dose, fractionation and technique, and outcomes include local control, toxicity and survival. Surgical trials compare operations, approaches such as open and minimally invasive surgery, and the extent of resection. These interventions depend on the skill of the team and on technology that changes quickly, so the results from older trials may not describe current practice. Blinding is usually impossible, and trials can be affected by differences in the care patients receive apart from the intervention. Reviews should record the technique, the center volume and the years of recruitment, and test whether results differ by era.

Supportive and palliative care trials use symptom scores, quality-of-life instruments and sometimes survival. Outcomes are often patient reported, with missing data due to death and decline, which makes the handling of missing data central. A review should state whether each trial analyzed outcomes among survivors only or used methods that account for death. Without that, quality-of-life results can look better than they are because the sickest patients are missing from later time points.

Common pitfalls we look for

  • Pooling hazard ratios from trials with different follow-up without noting that short follow-up favors early benefit.
  • Counting the same trial twice when the primary report and several updates or subgroup papers are all entered.
  • Treating response rate as proof of benefit without data on survival or quality of life.
  • Mixing lines of therapy and biomarker groups in one pooled estimate and then reporting it as the effect for all patients.
  • Ignoring industry sponsorship, since trials sponsored by manufacturers may differ in comparator choice and in the reporting of toxicity.
  • Reading a ranking from a network meta-analysis as a recommendation when the confidence intervals overlap widely.

Bias, certainty and reporting

Risk of bias is assessed with RoB 2 for randomized trials and ROBINS-I for non-randomized comparisons. Open-label design, subjective assessment of progression and imbalanced use of later therapy are the main concerns. Selective outcome reporting is documented in oncology, so registered outcomes should be compared with the publication. Certainty is rated with GRADE for each outcome, with attention to indirectness when the trial population differs from the question, and to imprecision when the number of events is low.

Reporting follows PRISMA 2020, with extensions for network meta-analysis and individual participant data where used. Protocols can be registered in PROSPERO. Toxicity should be reported along with efficacy, and quality-of-life data should not be omitted because they are harder to pool: a table of what each trial measured is more honest than silence.

How we support research projects in this area

Support

From a clinical question to a published review

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A structured question, eligibility criteria and an analysis plan, with registration prepared where appropriate.

  • Searching and extraction

    Search strategies for the relevant databases and registries, screening and data extraction, and risk-of-bias assessment by design.

  • Synthesis

    Pairwise, network, diagnostic accuracy, prognostic or dose-response analysis, with a GRADE assessment for each outcome.

  • Manuscript and submission

    Reporting-guideline checklists, the manuscript and the preparation of submission materials.

Get a quoteDescribe your question, study types and target journal.

Boundaries of this service

A review of cancer studies summarizes group-level evidence. It does not tell a person with cancer which treatment to choose, which depends on tumor features, other illness, preferences and the advice of the treating team. We do not provide treatment recommendations or advice on the care of any patient, or interpretation of an individual's results. Guideline development is a separate process with its own panels and methods.

Frequently asked questions

Why are hazard ratios used in oncology reviews?

Survival outcomes are times to events with censoring, and the hazard ratio summarizes the effect over follow-up. It is pooled on the log scale.

What if the proportional hazards assumption fails?

Report restricted mean survival time or landmark results, examine curves, and use individual patient data if available.

Are progression-free survival results reliable substitutes for survival?

It depends on the cancer type and treatment class. Trial-level analyses show variable association, so the setting should be stated.

Can single-arm trials be pooled?

They can be pooled to summarize response in treated patients, but not to estimate a treatment effect against a comparator.

How do I handle crossover after progression?

Record how each trial handled it, consider adjusted analyses where reported, and judge the certainty accordingly.

Do you give treatment advice?

No. The service provides research and evidence-synthesis support only.

References

  1. Early Breast Cancer Trialists' Collaborative Group (EBCTCG). Effects of chemotherapy and hormonal therapy for early breast cancer on recurrence and 15-year survival: an overview of the randomised trials. Lancet. 2005;365(9472):1687-1717.
  2. Parmar MKB, Torri V, Stewart L. Extracting summary statistics to perform meta-analyses of the published literature for survival endpoints. Stat Med. 1998;17(24):2815-2834.
  3. Tierney JF, Stewart LA, Ghersi D, Burdett S, Sydes MR. Practical methods for incorporating summary time-to-event data into meta-analysis. Trials. 2007;8:16.
  4. Prasad V, Kim C, Burotto M, Vandross A. The strength of association between surrogate end points and survival in oncology: a systematic review of trial-level meta-analyses. JAMA Intern Med. 2015;175(8):1389-1398.
  5. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  6. Royston P, Parmar MKB. Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome. BMC Med Res Methodol. 2013;13:152.
  7. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.