What the standardized mean difference is
When all studies measure an outcome on the same scale, such as systolic blood pressure in mmHg, the natural summary is the mean difference, which keeps its units and can be interpreted directly. When studies measure the same underlying construct with different instruments, for example several depression rating scales with different ranges, their mean differences cannot be pooled, because a five-point difference means different things on different scales.
The standardized mean difference solves this by dividing the difference between the group means by a standard deviation. The result has no units. It says how many standard deviations apart the two groups are, and that can be compared and pooled across scales. The price is that the quantity is a ratio, so it depends on the variability of the sample as well as on the size of the effect, a point that causes much of the confusion in interpretation.
Cohen's d and Hedges' g
Cohen's d is the difference in means divided by the pooled standard deviation of the two groups. The pooled standard deviation is the square root of the weighted average of the two variances, weighted by the degrees of freedom of each group.
Hedges' g multiplies d by a correction factor, usually written J, which is approximately 1 - 3 / (4 x (n1 + n2 - 2) - 1). In small samples, d overestimates the population effect slightly, and the factor reduces it. In large samples the factor is close to 1 and d and g are almost the same. Most meta-analysis software reports g, and the terms are sometimes used loosely, so check which one is being calculated.
The standard error of g is approximately the square root of (n1 + n2) / (n1 x n2) + g2 / (2 x (n1 + n2)). The first term is the sampling variability of the difference and the second reflects uncertainty in the standard deviation. Notice that the standard error depends on the effect itself, a fact that matters for publication bias tests.
A variant, Glass's delta, divides by the standard deviation of the control group alone. It is used when the intervention might change the variance, so that the control group provides a cleaner reference.
A worked example
Suppose a simulated trial reports mean scores of 24 (standard deviation 8, n = 50) in the intervention group and 20 (standard deviation 7, n = 50) in the control group.
- Difference in means: 4 points.
- Pooled standard deviation: 7.52.
- Cohen's d: 4 / 7.52 = 0.53.
- Correction factor J: 0.992.
- Hedges' g: 0.992 x 0.53 = 0.53.
- Standard error of g: 0.20, giving a 95 percent confidence interval of 0.13 to 0.93.
With 100 participants, d and g agree to two decimals. With a total of 20, the correction would reduce d by about four percent. The correction matters more when samples are small, which is when meta-analysts most often need it.
Pooling across scales
Three simulated trials measure depression on three different scales and report Hedges' g, with negative values favoring the intervention.
| Study | Sample sizes | Hedges' g | Standard error | 95% CI |
|---|---|---|---|---|
| Depression scale A | 40 / 40 | -0.62 | 0.23 | -1.07 to -0.17 |
| Depression scale B | 90 / 88 | -0.35 | 0.15 | -0.65 to -0.05 |
| Depression scale C | 25 / 27 | -0.48 | 0.28 | -1.03 to 0.07 |
Inverse-variance pooling under a fixed-effect model gives a pooled g of -0.44 (standard error 0.12). A real analysis would also consider heterogeneity and probably use a random-effects model. The example shows the basic mechanism: because each study's effect is expressed in standard deviations, the three scales can be combined.
This is only appropriate if the scales measure the same construct in the same direction. A review that combines a depression scale with a measure of quality of life, because both are labeled outcomes of the intervention, creates a pooled number that no one can interpret. Before pooling, check that the instruments are valid for the construct, that higher scores mean the same thing, and that, where scores run in opposite directions, the sign of one has been reversed.
Interpreting a standardized mean difference
Cohen suggested labels of small (0.2), medium (0.5) and large (0.8) for effect sizes in behavioral science, and these labels are quoted widely. He intended them as rough benchmarks when nothing else was known, and he warned against using them rigidly. They are not clinical thresholds. A standardized difference of 0.3 may be trivial for one condition and important for another, and the same standardized value corresponds to different raw differences in different populations.
Part of the problem is that a standardized difference depends on the standard deviation. If one trial recruits a narrow range of patients and another a broad range, the same raw difference gives different standardized values. The comparison of standardized values between studies therefore mixes the size of the effect with the heterogeneity of the sample. The quantity has an appealing scale and an awkward meaning.
Some alternative ways of expressing the result are better suited to decision making:
- Re-expression on a familiar scale. Multiply the pooled SMD by a typical standard deviation from a well-known instrument. For example, a pooled g of -0.44 on a scale with a standard deviation of 10 points corresponds to a difference of about 4.4 points.
- Comparison with a minimal important difference. If the minimal important difference for the scale is known, compare the re-expressed effect with it.
- Conversion to an odds ratio. The log odds ratio is approximately the SMD multiplied by pi divided by the square root of 3, about 1.81. A standardized difference of 0.53 corresponds to an odds ratio of about 2.6, assuming normally distributed outcomes. The assumption should be stated.
- Proportion of patients improved. Relative to an assumed control event rate, express the effect as the proportion above a threshold.
The Cochrane Handbook discusses these options, and recommends that authors do not leave readers with a bare SMD when a more understandable alternative can be provided.
Assumptions and cautions
The standardized mean difference assumes that the outcome is approximately normally distributed within groups and that variances are similar in the two groups. Skewed data, such as costs or length of stay, and ceiling or floor effects can invalidate it. When trials report medians and ranges, the means and standard deviations must be estimated, and the estimation adds uncertainty that should be explored in a sensitivity analysis.
Other points of care include the following.
- Change scores and final scores. Studies may report the change from baseline or the final value. Their standard deviations differ, and standardizing by one and not the other makes the effects incomparable. Final values and change scores should be handled consistently, or analysed separately.
- Cluster and crossover designs. Standard errors need adjustment for clustering or for the paired structure. Using raw sample sizes overstates precision.
- Multiple arms and multiple outcomes. A multi-arm trial must not enter the analysis twice with a shared control group. Several outcomes from the same sample are correlated.
- Heterogeneity of scales. Similar constructs can be measured by instruments that differ in content, and the pooled value then mixes them.
- Small-study effects. Because the standard error of the SMD depends on the SMD, tests for funnel plot asymmetry can behave differently from tests with a standard error that is independent of the effect.
How to report it
State whether d or g was used, the formula for the standard error, the model for pooling, the handling of change and final scores, the direction of effect, and any rescaling or sign reversal. Present study-level values, sample sizes and confidence intervals in a forest plot, and give the pooled estimate with an interval and a prediction interval for a random-effects analysis. Add an interpretation that does not rely on Cohen's labels. A reader should be able to see what the pooled value means on a scale they know. GRADE guidance supports presenting the result in terms of a familiar instrument or a relative measure in a summary of findings table.
Calculating an SMD from what trials report
Trial reports rarely give the pooled standard deviation. Extraction usually starts from whatever is available, and the route matters.
- Means, standard deviations and sample sizes. The ideal case, and the formulas above apply directly.
- Standard errors or confidence intervals for each mean. Convert to standard deviations by multiplying the standard error by the square root of the sample size. For a confidence interval from a t distribution in a small sample, use the appropriate multiplier and not 1.96.
- A t statistic or p value for the difference. The SMD can be recovered from the t value and the group sizes. An exact p value is far more useful than a statement of significance.
- Medians and ranges or interquartile ranges. Methods exist to estimate a mean and standard deviation, but they assume approximately symmetric data, and results should be checked by excluding such studies in a sensitivity analysis.
- Adjusted means from a regression or covariance model. The standard deviation may be smaller than the unadjusted one, so mixing adjusted and unadjusted effects across studies needs thought.
Record not only the number but where it came from, so that another person can check it. Many reviews double-extract effect data because errors in this step are common and consequential.
How we can help
We can choose the appropriate effect measure, calculate d and g from reported statistics including medians and confidence intervals, handle change and final scores, pool across scales, re-express results on a familiar scale and prepare summary of findings tables. [OWNER VERIFICATION REQUIRED] The relevant services are meta-analysis and statistical analysis.
Frequently asked questions
What is the difference between Cohen's d and Hedges' g?
Hedges' g applies a small-sample correction to d. The two are nearly identical in large samples, and g is the usual choice in meta-analysis.
When should I use an SMD instead of a mean difference?
When studies measure the same construct on different scales. If all studies use the same scale, the mean difference is easier to interpret.
Are 0.2, 0.5 and 0.8 reliable thresholds?
They are conventions from behavioral science and not clinical thresholds. Interpret an SMD against a minimal important difference or re-express it on a familiar scale.
How do I convert an SMD to an odds ratio?
The log odds ratio is about the SMD multiplied by pi divided by the square root of 3, which is about 1.81. It assumes normally distributed outcomes.
Can I pool change scores with final scores?
Not as if they were the same. Their standard deviations differ, so handle them consistently or analyse them separately.
Why does the SMD depend on the sample's variability?
Because it divides by a standard deviation. A narrower sample gives a larger SMD for the same raw difference.
References
- Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Lawrence Erlbaum Associates; 1988.
- Hedges LV. Distribution theory for Glass's estimator of effect size and related estimators. J Educ Stat. 1981;6(2):107-128.
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
- Deeks JJ, Higgins JPT, Altman DG, McKenzie JE, Veroniki AA. Chapter 10: Analysing data and undertaking meta-analyses. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Chinn S. A simple method for converting an odds ratio to effect size for use in meta-analysis. Stat Med. 2000;19(22):3127-3131.
- Wan X, Wang W, Liu J, Tong T. Estimating the sample mean and standard deviation from the sample size, median, range and/or interquartile range. BMC Med Res Methodol. 2014;14:135.
- Guyatt GH, Thorlund K, Oxman AD, et al. GRADE guidelines: 13. Preparing summary of findings tables and evidence profiles: continuous outcomes. J Clin Epidemiol. 2013;66(2):173-183.