Method

Prognostic meta-analysis

Prognostic meta-analysis synthesizes evidence on what predicts the future course of a condition: single prognostic factors and multivariable prediction models. It is harder than meta-analysis of treatment trials, because the studies are observational, adjust for different variables, and report results in varied ways.

What prognosis research asks

Prognosis research studies the course and outcome of health conditions and the factors that predict them. It asks four kinds of question. What is the overall prognosis of a condition, in terms of how many people experience an outcome over time? Which factors are associated with a better or worse outcome, independent of others? How well can a combination of factors predict an individual's outcome, in the form of a prediction model? And which factors modify the response to treatment, so that a marker could guide therapy? Each leads to different studies and different synthesis methods. Prognosis is distinct from diagnosis, which concerns present status, and from treatment effects, which concern the impact of an intervention.

The research is overwhelmingly observational, so synthesis must deal with confounding, with differences between studies in what they adjust for, and with selective reporting of results. The methods have developed substantially in recent years, with dedicated tools and guidance. This page outlines them. The general principles are in meta-analysis.

Meta-analysis of prognostic factors

A prognostic factor study estimates the association between a factor and an outcome, usually as a hazard ratio for time-to-event outcomes, or an odds or risk ratio for binary ones. The central problem is that the estimates are not comparable. One study may report an unadjusted association, another an association adjusted for age and sex, a third for a large set of clinical variables, and the amount of adjustment changes the estimate. The factor may be measured as a continuous variable in one study and as high versus low at different cut-points in another. The outcome may be defined differently, and follow-up lengths differ.

The recommended approach is to synthesize unadjusted and adjusted estimates separately, and where possible to define a minimum set of adjustment variables in advance and to pool estimates that meet it. Studies with different cut-points can be pooled by converting results to a common scale, such as the estimate per standard deviation, when the necessary information is available. Random-effects models are standard, with attention to heterogeneity, which is typically large. Because prognostic factor estimates are particularly vulnerable to selective reporting, sensitivity analyses and checks for small-study effects matter. See the guide on pooling hazard ratios.

Extracting hazard ratios

Many prognostic studies report survival outcomes, and the hazard ratio and its confidence interval are the preferred results. When they are reported, the log hazard ratio and its standard error, derived from the interval, are pooled with the inverse-variance method. When they are not, methods exist to estimate them from other published information, such as the log-rank statistic and its p-value, the numbers of events and the sample sizes, or from published Kaplan-Meier curves by reconstructing approximate individual data. These estimates are less reliable than reported ones and are flagged, and a sensitivity analysis excludes studies whose estimates were derived. Multivariable results must come from comparable models. A hazard ratio from a model that includes many variables is not equivalent to one from a simpler model, and the interpretation of the pooled value must make clear what the adjustment was.

Meta-analysis of prediction models

A prediction model combines several predictors to estimate an individual's probability of an outcome. A model is developed in one dataset and should then be validated in others. The performance of a model in a new dataset is described by discrimination, which is its ability to separate people who will and will not have the outcome, usually reported as the concordance statistic or area under the curve, and calibration, which is the agreement between predicted and observed risks, summarized by the calibration slope and the calibration-in-the-large. Calibration is at least as important as discrimination, since a model that ranks people well but systematically over- or under-predicts risk can cause harm when used for decisions.

A meta-analysis of prediction model studies typically pools the performance of a specific model across external validation studies. It estimates the average discrimination and calibration and, importantly, how much they vary across settings and populations, using random-effects models, often on transformed scales such as the logit of the concordance statistic, and prediction intervals show the range of performance to expect in a new setting. Heterogeneity in performance is common and expected, because case mix and baseline risk differ. A model whose performance varies much across settings may need local recalibration before use. See the guide in the references by Debray and colleagues.

Appraisal and data extraction tools

QUIPS
The Quality In Prognosis Studies tool, for prognostic factor studies, examines six domains: participation, attrition, prognostic factor measurement, outcome measurement, confounding and statistical analysis and reporting.
PROBAST
The Prediction model Risk Of Bias ASsessment Tool examines participants, predictors, outcome and analysis, and assesses applicability to the review question. Analysis problems, such as small sample size relative to the number of candidate predictors and inadequate handling of missing data, are common reasons for high risk.
CHARMS
A checklist for critical appraisal and data extraction in reviews of prediction modeling studies, guiding the items to extract about source of data, participants, outcome, candidate predictors, sample size, missing data, model development and performance.
REMARK and TRIPOD
Reporting guidelines for prognostic marker studies and for prediction model studies, which review authors can use as markers of reporting quality.

See risk-of-bias tools.

Individual participant data

Prognosis is an area in which individual participant data meta-analysis is especially valuable. With participant-level data, the analyst can adjust for the same variables in every study, use the same cut-points and outcome definitions, check assumptions such as proportional hazards, examine nonlinear relations and interactions, and validate prediction models in many datasets with the same procedures. It also reduces the selective reporting that affects published results. The cost is the effort of obtaining and harmonizing data. Where participant-level data cannot be obtained, aggregate data from published reports is used with the limitations described above. See individual participant data meta-analysis.

Bias, heterogeneity and certainty

Several threats are prominent. Selective reporting is a serious problem because prognostic studies often explore many factors and cut-points, and report those that are significant. Publication bias affects the evidence for factors that appear predictive. Confounding is rarely resolved, since adjusting for measured variables does not remove unmeasured confounding. Many studies are small, with few outcome events relative to candidate predictors, which leads to overfitting and inflated associations. The heterogeneity is nearly always large, owing to differences in populations, treatments and follow-up. For these reasons the certainty of evidence from prognostic reviews is often low, and approaches have been adapted to rate it. Funnel plots and tests for small-study effects are used when enough studies exist, with caution. Sensitivity analyses can restrict the analysis to studies at lower risk of bias or to those with adequate sample size.

Overall prognosis

The simplest prognostic question is the average course of a condition: what proportion are alive or free of recurrence after one year, five years, or longer. Studies report survival curves, event rates at fixed time points or median survival, and these have to be brought to a common form. Pooling event proportions at a fixed time point is straightforward when the time points match and follow-up is complete, and the methods are those for proportions. Where follow-up differs or participants are censored, pooling raw proportions is misleading, and the analysis uses rates, hazard-based summaries or survival probabilities reconstructed from curves. Median survival is a poor input to meta-analysis because its variance is rarely reported and it depends on the length of follow-up. A review of overall prognosis should state the time frame and the way censoring was handled, because they determine what the pooled number means.

Treatment-effect modifiers

Some prognostic research asks not whether a factor predicts outcome, but whether it predicts who benefits from a treatment, which is a question of interaction between the factor and the treatment. Such questions need trial data, and the right analysis estimates the interaction within each trial and pools the interactions, because comparing treatment effects across studies, or across subgroups defined by the factor, mixes within-trial and between-trial information and is open to ecological bias. Individual participant data are strongly preferred. Claims that a marker predicts treatment benefit are among the most commonly overturned in the literature, and subgroup findings from single trials should be treated as hypotheses. Where a review addresses such a question, it should state in advance the interactions of interest, report them with their intervals, and apply the cautions described for meta-regression.

Validation and updating of prediction models

Most published prediction models are developed and evaluated in the same data, which overstates their performance. A model is credible only after external validation in data that were not used to develop it, ideally from different places and times. A review of a model collects the validation studies and examines how it performs in each. When calibration is poor in new settings, which is common because baseline risk differs, the model can be updated: the intercept can be re-estimated for local risk, the slope adjusted, or the coefficients refit, and the updated model should itself be validated. Reviews also compare competing models for the same outcome, which needs head-to-head validation in the same data to be fair. These steps are what separates a model that is merely published from one that is ready to be tested in practice, and whether the model changes decisions and outcomes is a further question that accuracy alone cannot answer.

Planning a review

  1. State the question precisely

    The population, the factor or model, the outcome and the time frame are defined, and the purpose of the review is stated: overall prognosis, a factor, a model, or a treatment modifier.

  2. Plan the search

    Prognosis studies are poorly indexed and design filters perform unevenly, so searches rely on concept combinations, with registries and conference abstracts considered.

  3. Define the adjustment strategy

    The minimum set of adjustment variables, the preferred cut-points and the handling of unadjusted estimates are specified in advance.

  4. Extract and appraise

    CHARMS guides extraction and QUIPS or PROBAST the appraisal.

  5. Synthesize and report

    Estimates are pooled with random effects, heterogeneity is explored, and the review is reported with PRISMA 2020 and with the guidance specific to prognosis.

Limitations

Pooled prognostic estimates are vulnerable to differences in adjustment, definition and design that cannot always be corrected. The studies are observational, so associations are not causal. Prediction models can fail when applied to new settings, and a pooled performance measure hides this unless the heterogeneity is reported. Many models are developed in small datasets and have not been externally validated. The review does not decide whether a model should be used in practice, which requires evidence of impact on decisions and outcomes. Results are not advice about the care of an individual patient.

Prognostic research informs risk estimation. It does not provide clinical advice for individuals.

How we can help

Support

Support for prognostic meta-analysis

The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.

Get a quoteTell us your question, the number of studies and your target journal.

Frequently asked questions

What is the difference between a prognostic factor and a prediction model?

A prognostic factor is a single variable associated with the outcome. A prediction model combines several predictors to estimate an individual's risk.

Should I pool adjusted and unadjusted estimates together?

Usually not. They estimate different things. Pool them separately, and define in advance a minimum set of adjustment variables for the adjusted analysis.

How is the performance of a prediction model meta-analyzed?

The discrimination and calibration of a specific model across external validation studies are pooled with random-effects models, with prediction intervals showing how much performance varies between settings.

What tools assess the quality of prognostic studies?

QUIPS for prognostic factor studies and PROBAST for prediction model studies, with CHARMS for data extraction in prediction model reviews.

Why is individual participant data so useful here?

It allows consistent adjustment, cut-points and outcome definitions across studies, checks of assumptions, and uniform validation of models, and it reduces selective reporting.

Is a high pooled hazard ratio proof that a factor is causal?

No. The studies are observational and the association may reflect confounding. Prognostic factors are markers of risk and not necessarily causes.

References

  1. Riley RD, Moons KGM, Snell KIE, et al. A guide to systematic review and meta-analysis of prognostic factor studies. BMJ. 2019;364:k4597.
  2. Debray TPA, Damen JAAG, Snell KIE, et al. A guide to systematic review and meta-analysis of prediction model performance. BMJ. 2017;356:i6460.
  3. Moons KGM, de Groot JAH, Bouwmeester W, et al. Critical appraisal and data extraction for systematic reviews of prediction modelling studies: the CHARMS checklist. PLoS Med. 2014;11(10):e1001744.
  4. Hayden JA, van der Windt DA, Cartwright JL, Cote P, Bombardier C. Assessing bias in studies of prognostic factors. Ann Intern Med. 2013;158(4):280-286.
  5. Wolff RF, Moons KGM, Riley RD, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58.
  6. Moons KGM, Wolff RF, Riley RD, et al. PROBAST: a tool to assess risk of bias and applicability of prediction model studies: explanation and elaboration. Ann Intern Med. 2019;170(1):W1-W33.
  7. Riley RD, Hayden JA, Steyerberg EW, et al. Prognosis Research Strategy (PROGRESS) 2: prognostic factor research. PLoS Med. 2013;10(2):e1001380.
  8. Steyerberg EW, Moons KGM, van der Windt DA, et al. Prognosis Research Strategy (PROGRESS) 3: prognostic model research. PLoS Med. 2013;10(2):e1001381.
  9. Collins GS, Reitsma JB, Altman DG, Moons KGM. Transparent reporting of a multivariable prediction model for individual prognosis or diagnosis (TRIPOD): the TRIPOD statement. BMJ. 2015;350:g7594.
  10. Parmar MKB, Torri V, Stewart L. Extracting summary statistics to perform meta-analyses of the published literature for survival endpoints. Stat Med. 1998;17(24):2815-2834.
  11. Tierney JF, Stewart LA, Ghersi D, Burdett S, Sydes MR. Practical methods for incorporating summary time-to-event data into meta-analysis. Trials. 2007;8:16.
  12. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.