What a hazard ratio is
A hazard is the instantaneous rate at which an event, such as death or disease progression, occurs among people who have not yet had it. The hazard ratio compares the hazard in two groups. A hazard ratio of 0.80 means that, at any point in follow-up, people in the intervention group have an event at 80 percent of the rate of people in the control group, among those still event-free. It is a relative measure of the event rate over time.
Time-to-event outcomes need a special measure because follow-up differs between participants. Some have the event, some are lost to follow-up, and some are still event-free when the study ends, which is called censoring. A simple proportion of people with the event ignores when it happened and who was observed for how long. The hazard ratio uses that information, and it is estimated from a Cox proportional hazards model or from a logrank analysis.
It is easy to misread. A hazard ratio is not a risk ratio and does not say that the risk of the event by a given time is reduced by a given percentage. It describes a ratio of rates. Under the assumption that the ratio is constant over time, it can be converted to survival at a chosen time, as shown below, but the conversion needs a baseline.
Pooling on the log scale
Hazard ratios are pooled on the log scale, for the same reasons as other ratio measures: the log hazard ratio is approximately normally distributed, the scale is symmetric around no effect (log 1 = 0), and a hazard ratio of 2 and one of 0.5 are equidistant from the null. Each study contributes its log hazard ratio and the standard error of that quantity. The weights are the inverse of the variance, and the pooled log hazard ratio is the weighted average. Exponentiating the result and its confidence limits returns to the hazard ratio scale.
Most reports give a hazard ratio with a 95 percent confidence interval. The standard error of the log hazard ratio is then (log of the upper limit minus log of the lower limit) divided by 3.92, where 3.92 is twice 1.96. The example below uses this.
| Trial | Hazard ratio | 95% CI | log HR | SE of log HR | Weight |
|---|---|---|---|---|---|
| Trial A | 0.75 | 0.60 to 0.94 | -0.288 | 0.115 | 76.2 |
| Trial B | 0.82 | 0.66 to 1.02 | -0.198 | 0.111 | 81.1 |
| Trial C | 0.68 | 0.50 to 0.92 | -0.386 | 0.156 | 41.3 |
| Trial D | 0.90 | 0.71 to 1.14 | -0.105 | 0.121 | 68.5 |
The weighted mean of the four log hazard ratios is -0.229, with a standard error of 0.061. Exponentiating gives a pooled hazard ratio of 0.80 (95 percent CI 0.71 to 0.90) under a fixed-effect model. Three of the four trials have intervals that include 1, and the pooled interval does not, which is the usual benefit of pooling. A real review would also examine heterogeneity and normally use a random-effects model. The trials are simulated.
When the hazard ratio is not reported
Many trials do not report a hazard ratio, or report one that is unusable, for instance an unadjusted figure from a model that is not comparable with the others. Parmar and colleagues described methods to extract the log hazard ratio and its variance from other published information, and Tierney and colleagues set out a practical guide to doing it. The options depend on what is reported.
- Hazard ratio with confidence interval. Use directly, as above.
- Hazard ratio with an exact p value. The standard error can be calculated from the p value and the estimate.
- Logrank statistic or observed and expected events. The log hazard ratio is approximately (observed minus expected events) divided by the variance, and the standard error is the reciprocal of the square root of the variance.
- Kaplan-Meier curves with numbers at risk. Individual patient data can be approximately reconstructed from digitized curves and the numbers at risk, using an algorithm published by Guyot and colleagues, and a hazard ratio can then be estimated. This is the least certain route and should be flagged.
- Median survival times only. Not enough for a valid hazard ratio without strong assumptions, and usually best avoided.
Reconstructed estimates carry extra uncertainty, and the method assumes that the published curve is faithful and the number at risk is given at enough time points. A sensitivity analysis excluding reconstructed estimates is good practice. Two extractors should do the work independently, as digitizing curves is prone to small differences.
The proportional hazards assumption
A single hazard ratio describes a whole follow-up period only if the hazards are proportional, meaning the ratio is constant over time. This is not guaranteed. Immunotherapy trials in oncology, for example, often show curves that overlap early and separate later, a delayed effect, and some surgical comparisons show an early excess of events followed by a benefit. The curves can also cross. In these situations the hazard ratio depends on the length of follow-up, and a single number can mislead.
What can be done in a review:
- Inspect the curves. Look for crossing or a delayed separation, where the original papers allow it.
- Look for tests and plots in the trials. Some report a test of proportionality or a plot of the log cumulative hazard.
- Consider alternatives. Restricted mean survival time, which is the average event-free time up to a chosen horizon, does not require proportional hazards, and the difference or ratio of restricted mean survival times is easier to interpret. Landmark analyses and milestone survival at chosen time points are other options.
- Explore the effect of follow-up. Meta-regression on median follow-up, or a subgroup analysis by follow-up length, can show whether the hazard ratio changes.
- Be careful with the wording. If non-proportionality is plausible, describe the pooled hazard ratio as an average over the observed follow-up and not as a constant effect.
The need for these checks is a limitation of the whole approach. Pooling several hazard ratios from trials with different follow-up periods can combine different parts of the curves, and heterogeneity may reflect that.
From a hazard ratio to an absolute effect
Under proportional hazards, survival in the intervention group at any time t is the control survival raised to the power of the hazard ratio: S1(t) = S0(t) raised to the power HR. Suppose control survival at two years is 60 percent and the pooled hazard ratio is 0.80. Then survival with the intervention is 0.60 raised to 0.80, which is 0.666, or about 66.6 percent. The absolute difference is about 6.6 percentage points at two years.
The same hazard ratio gives different absolute gains at different baselines or time points. With very poor survival, the absolute gain is limited by how many are still alive to benefit, and with very good survival, there are few events to prevent. Reporting a hazard ratio alone therefore leaves out the figure patients and clinicians need. Summary of findings tables in GRADE usually give the absolute difference at a stated time point, with the assumed control risk and its source.
Heterogeneity and model choice
Differences between trials in populations, treatment doses, follow-up and the definition of the event are common in survival data. A random-effects model allows the true hazard ratio to vary between trials and gives wider intervals when it does. Report I-squared and tau-squared with the usual cautions, and add a prediction interval when there are enough trials. Pre-planned subgroup analyses or meta-regression can explore whether the effect depends on a characteristic, for example a biomarker status, the line of therapy or the median follow-up.
Subgroup hazard ratios from within a trial are a source of overinterpretation, because they come from smaller samples and from many comparisons. Interaction tests are more informative than a comparison of the significance of the separate subgroups.
Common pitfalls
- Mixing adjusted and unadjusted hazard ratios. In randomized trials the unadjusted figure is usually preferred for consistency. In observational studies, use the most fully adjusted estimate, and record the covariates.
- Using the wrong event. Overall survival, progression-free survival, disease-free survival and event-free survival are different outcomes. Pooling them together is rarely meaningful, and trial-level correlations between surrogate endpoints and survival vary in strength.
- Treating a hazard ratio as a risk ratio. The statement that treatment cuts the risk of death by 20 percent is not correct for a hazard ratio of 0.80.
- Confusing the direction. Check which group is the reference in each trial and invert where required, taking the reciprocal of the hazard ratio and swapping the limits.
- Competing risks. When another event prevents the one of interest, the standard hazard may not answer the question, and subdistribution hazards need to be considered with care.
- Ignoring multi-arm trials. The shared control arm must not be counted twice.
- Immature data. Hazard ratios from early analyses with few events are unstable and can change materially on later follow-up.
Reporting
State the event definition, the source of each hazard ratio (reported, derived from other statistics, or reconstructed from curves), the model, the direction, and the treatment of multi-arm trials. Show a forest plot with hazard ratios and intervals on a logarithmic axis. Report the assessment of proportional hazards and the follow-up of each trial. Present absolute effects at one or more stated time points, together with the assumptions behind them. The reporting of a systematic review that includes such outcomes follows PRISMA 2020, and the certainty of evidence is rated with GRADE.
How we can help
We can extract hazard ratios and derive missing ones, reconstruct data from published curves with a documented method, assess proportional hazards, run fixed-effect and random-effects analyses, and calculate absolute effects for summary of findings tables. [OWNER VERIFICATION REQUIRED] The relevant services are meta-analysis and statistical analysis.
Frequently asked questions
Why are hazard ratios pooled on the log scale?
Because the log hazard ratio is approximately normal and symmetric around no effect. Studies are combined on that scale and the result is exponentiated.
How do I get the standard error from a confidence interval?
Subtract the log of the lower limit from the log of the upper limit and divide by 3.92 for a 95 percent interval.
What if a trial does not report a hazard ratio?
Derive it from a p value, logrank statistics or observed and expected events, or reconstruct data from Kaplan-Meier curves. Flag and test the effect of such estimates.
What is the proportional hazards assumption?
That the hazard ratio is constant over follow-up. If curves cross or separate late, a single hazard ratio can mislead, and restricted mean survival time is an alternative.
Is a hazard ratio of 0.80 a 20 percent reduction in risk?
Not strictly. It is a 20 percent lower event rate at any time, among those still event-free. The effect on risk by a given time depends on the baseline.
Can I pool overall survival with progression-free survival?
Generally not. They are different outcomes with different meanings and should be analysed separately.
References
- Parmar MKB, Torri V, Stewart L. Extracting summary statistics to perform meta-analyses of the published literature for survival endpoints. Stat Med. 1998;17(24):2815-2834.
- Tierney JF, Stewart LA, Ghersi D, Burdett S, Sydes MR. Practical methods for incorporating summary time-to-event data into meta-analysis. Trials. 2007;8:16.
- Guyot P, Ades AE, Ouwens MJNM, Welton NJ. Enhanced secondary analysis of survival data: reconstructing the data from published Kaplan-Meier survival curves. BMC Med Res Methodol. 2012;12:9.
- Royston P, Parmar MKB. Restricted mean survival time: an alternative to the hazard ratio for the design and analysis of randomized trials with a time-to-event outcome. BMC Med Res Methodol. 2013;13:152.
- Hernan MA. The hazards of hazard ratios. Epidemiology. 2010;21(1):13-15.
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA. Chapter 10: Analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Prasad V, Kim C, Burotto M, Vandross A. The strength of association between surrogate end points and survival in oncology: a systematic review of trial-level meta-analyses. JAMA Intern Med. 2015;175(8):1389-1398.