What an effect size is
A p-value says whether a result is compatible with no effect. It does not say how large the effect is, or whether it matters. An effect size does: it is a quantity, such as a difference in means or a ratio of risks, that measures the magnitude of an effect in a form that can be interpreted and compared. Meta-analysis combines effect sizes, and each study must therefore provide one, on a common scale, together with a measure of its precision. The choice of effect size shapes both the statistics and the interpretation.
Three properties are worth looking for. The measure should suit the type of data, so that it can be computed from what studies report. It should behave well statistically, with a variance that can be estimated and a distribution that is close to normal on the scale used for pooling. And it should be interpretable, by the readers who will use the result, in terms they understand. These aims sometimes conflict, and the choice is a compromise. The general approach is in meta-analysis.
Binary outcomes: risk ratio, odds ratio and risk difference
For outcomes that either happen or do not, three measures are common. The risk ratio is the risk in the intervention group divided by the risk in the control group. The odds ratio is the ratio of the odds, where odds are the number with the outcome divided by the number without. The risk difference is the absolute difference between the risks.
A small example shows how they relate. Suppose 30 of 100 people in the intervention group have the outcome, against 20 of 100 in the control group. The risks are 0.30 and 0.20. The risk ratio is 1.50, the odds ratio is (30/70) divided by (20/80), which is 1.71, and the risk difference is 0.10. The odds ratio is further from 1 than the risk ratio, and the gap widens as the outcome becomes more common. When the outcome is rare, the two are nearly equal. The figures are invented for illustration.
The choice follows several considerations. Relative measures, the risk ratio and odds ratio, tend to be more consistent across studies with different baseline risks than the risk difference, so they are usually pooled and then converted to absolute effects for particular baseline risks. The odds ratio has convenient statistical properties and is what logistic regression and case-control studies produce, but it is hard to interpret and is often misread as a risk ratio. The risk ratio is easier to understand but is bounded by the baseline risk. The guide on odds ratio versus risk ratio explains the difference further.
Continuous outcomes: mean difference and standardized mean difference
When studies measure an outcome on the same scale, the mean difference, the average in one group minus the average in the other, is the natural choice. It keeps the units, so a reader can see that the result is, say, 4 points on a 100-point scale, and relate it to a known threshold for importance.
When studies measure the same construct on different scales, the means cannot be combined directly, and the standardized mean difference is used. It divides the difference in means by a pooled standard deviation, giving a result in standard deviation units. In an example, a treatment group has a mean of 24 and standard deviation of 8, and the control group a mean of 20 and standard deviation of 7, with 50 people in each. The pooled standard deviation is 7.52 and the standardized difference, usually called Cohen's d, is 0.53. A correction factor of 0.992 for small samples gives Hedges' g of 0.53, with a standard error of about 0.20. The numbers are invented for illustration.
The standardized measure has costs. It depends on the spread of outcomes in each study, so a study with a narrow sample will show a larger standardized effect for the same raw difference. It is hard to interpret clinically, and the conventional labels of small, medium and large are rough guides at best. Where possible, the pooled result is converted back to a familiar scale. See standardized mean differences.
Time-to-event outcomes: the hazard ratio
For outcomes such as survival or time to relapse, in which participants are followed for different lengths of time and some are not observed to the event, the hazard ratio is the standard measure. It compares the instantaneous rate of the event in the two groups, and is estimated by Cox models and similar methods. A hazard ratio of 0.75 means that, at any time, the rate of events in the treatment group is 75 percent of that in the control group. It is pooled on the log scale using its standard error, which is derived from the reported confidence interval when it is not stated directly.
The hazard ratio assumes the ratio is constant over time, which may not hold. If survival curves cross, or the effect appears late, a single ratio can mislead, and other summaries, such as differences in restricted mean survival time, are sometimes preferred. Hazard ratios from trials and from observational studies differ in meaning, since adjusted observational estimates depend on the variables in the model. See pooling hazard ratios.
Associations: correlation
Where studies report associations between two continuous variables, the correlation coefficient is the usual measure. Because its distribution is skewed, and its variance depends on its value, correlations are transformed before pooling with Fisher's z transformation, which is half the natural log of (1 plus r) divided by (1 minus r). The variance of z depends only on the sample size, being 1 divided by the sample size minus 3. With a correlation of 0.30 in 100 people, z is 0.310 with a standard error of 0.102, and the pooled value is transformed back for presentation. Corrections for measurement unreliability and range restriction are sometimes applied in psychology and management to estimate the relationship between underlying constructs, a practice with its own assumptions. See correlation-based meta-analysis.
Choosing a measure
| Data | Usual measure | Pooled on | Main caution |
|---|---|---|---|
| Binary events | Risk ratio or odds ratio | Log scale | Odds ratios exaggerate risk ratios for common outcomes |
| Binary events, absolute effect | Risk difference | Original scale | Varies with baseline risk, so often inconsistent |
| Continuous, same scale | Mean difference | Original scale | Requires identical scales and direction |
| Continuous, different scales | Standardized mean difference (Hedges' g) | Standard deviation units | Depends on sample variability; hard to interpret |
| Time to event | Hazard ratio | Log scale | Proportional hazards assumption |
| Association | Correlation | Fisher's z | Range restriction and unreliable measurement |
| Single group | Proportion or rate | Logit or log scale | Extreme heterogeneity; unrepresentative samples |
Converting between measures
Studies often report results in forms that differ from the one chosen for the analysis, and conversions are common. A standardized mean difference can be converted to an odds ratio by an approximate formula in which the log odds ratio equals the standardized difference multiplied by pi divided by the square root of 3, about 1.81, a conversion used to put continuous and binary outcomes on one scale. For example, a standardized difference of 0.5 corresponds to a log odds ratio of about 0.91 and an odds ratio of about 2.5. A correlation can be converted to a standardized mean difference and back. Standard errors can be obtained from confidence intervals, and standard deviations from standard errors and sample sizes. Such conversions rest on assumptions, such as normally distributed outcomes or a logistic distribution, and they add uncertainty. They are applied with care, documented in the extraction dataset, and examined in sensitivity analysis. The service for data extraction and coding handles them.
From relative to absolute effects
Relative measures are good for pooling and poor for decisions, because the same relative effect means different things at different baseline risks. A risk ratio of 0.80 reduces a baseline risk of 10 percent to 8 percent, which is 20 fewer events per 1,000 people, but it reduces a risk of 1 percent to 0.8 percent, which is only 2 fewer per 1,000. A report should therefore convert the pooled relative effect into absolute effects for one or more stated baseline risks, drawn from the control groups of the studies or from a population of interest, and present them with the number needed to treat where it is helpful. This conversion is an essential part of the summary of findings table described in GRADE.
Effect size and importance
A statistically significant effect can be too small to matter, and a non-significant one can conceal an important effect in an imprecise estimate. For that reason, interpretation compares the effect size with a threshold of importance. For clinical outcomes the minimal important difference, the smallest change that patients perceive as beneficial, is used where it has been established for the measure. For other outcomes, a context-specific judgment is needed, informed by the field. The confidence interval is compared with the threshold: if it lies entirely above it, the effect is important with confidence, if entirely below, it is unimportant, and if it straddles the threshold, the evidence cannot say. This gives a more useful reading than the p-value alone, and it is the basis of the imprecision judgment in GRADE. Authors should state the threshold they used, and why, and avoid describing effects as small, medium or large without a reference.
Computing effect sizes in practice
Effect sizes and their variances can be computed by hand for a few studies, and in practice are computed by software, in the same scripts that do the meta-analysis, so that every step is reproducible. In R, the metafor package offers a function that takes raw counts, means and standard deviations, or correlations, and returns effect sizes and sampling variances for a wide choice of measures. Other programs have equivalent features. The analyst specifies the measure, checks the direction, and inspects the resulting values for implausible ones, such as a very large standardized difference caused by a misplaced decimal. The effect sizes, standard errors and the conversions applied are kept as part of the extraction dataset, so that a reader can check how each was obtained. This is also where the unit of analysis is dealt with, since the software computes the variance for the data as entered and cannot tell if the design was clustered.
Common problems
- Mixing measures that cannot be combined, such as odds ratios and risk ratios, or change scores and final values.
- Reversing the direction of a scale, so that a higher score is better in some studies and worse in others, without recoding.
- Using the wrong standard deviation, for instance a standard error as if it were a standard deviation.
- Ignoring the unit of analysis, as in cluster-randomized or crossover trials, so that the variance is wrong.
- Interpreting a standardized difference with fixed labels, as if small, medium and large had a universal meaning.
- Counting adjusted and unadjusted estimates together from observational studies.
Support
Choosing and computing effect sizes, with documented conversions, is part of the data extraction and meta-analysis services.
Get a quoteTell us your outcomes and how studies report them.
Frequently asked questions
Which effect size should I use for a binary outcome?
The risk ratio or the odds ratio are usual for pooling, because they tend to be more consistent across studies than the risk difference. The odds ratio overstates the risk ratio when the outcome is common, so state which is used and convert to absolute effects for interpretation.
What is the difference between Cohen's d and Hedges' g?
Both are standardized mean differences. Hedges' g includes a correction for the upward bias of d in small samples, and it is preferred in meta-analysis.
Why are ratios analyzed on the log scale?
Because the sampling distribution of a log ratio is closer to normal and its variance is simpler. Results are transformed back for reporting.
Why transform correlations with Fisher's z?
Because correlations have a skewed distribution and a variance that depends on their value, and z has an approximately normal distribution with a variance that depends only on the sample size.
Can I combine mean differences and standardized mean differences?
Not directly. Studies on the same scale can be pooled as mean differences, and those on different scales as standardized mean differences. Mixed cases need conversion, with its assumptions stated.
How do I turn a pooled risk ratio into something clinicians can use?
Compute the absolute effect for a stated baseline risk, and, if helpful, the number needed to treat, and present them in a summary of findings table.
References
- Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
- Hedges LV. Distribution theory for Glass's estimator of effect size and related estimators. J Educ Stat. 1981;6(2):107-128.
- Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Lawrence Erlbaum; 1988.
- Cooper H, Hedges LV, Valentine JC, editors. The Handbook of Research Synthesis and Meta-Analysis. 3rd ed. Russell Sage Foundation; 2019.
- Chinn S. A simple method for converting an odds ratio to effect size for use in meta-analysis. Stat Med. 2000;19(22):3127-3131.
- Zhang J, Yu KF. What's the relative risk? A method of correcting the odds ratio in cohort studies of common outcomes. JAMA. 1998;280(19):1690-1691.
- Hunter JE, Schmidt FL. Methods of Meta-Analysis: Correcting Error and Bias in Research Findings. 3rd ed. Sage; 2015.
- Parmar MKB, Torri V, Stewart L. Extracting summary statistics to perform meta-analyses of the published literature for survival endpoints. Stat Med. 1998;17(24):2815-2834.