What a forest plot is
A forest plot is the standard display of a meta-analysis. It gives one row for each study, showing the study's effect estimate and its uncertainty, and a final row for the pooled result. The name is said to come from the resemblance of the lines to a forest of trees. The format lets a reader see at a glance how large the effect was in each study, how precisely it was estimated, how much the studies agree, and what the overall result is, without reading any numbers. Because it shows the data behind the summary, it makes a meta-analysis open to inspection, which is its main value.
The same layout is used for any effect measure, whether risk ratios, odds ratios, hazard ratios, mean differences or standardized mean differences, and for subgroup analyses, in which the studies are grouped and each group has its own pooled result. The figure below is an annotated example built from simulated data, not from real research. The page on meta-analysis explains how the pooled result is calculated.
An annotated example
- 1. Study label. The first author and year, or an identifier, in the usual form.
- 2. Square. The study's point estimate. Its size is proportional to the weight the study carries in the analysis, so larger squares mean more influence.
- 3. Horizontal line. The 95 percent confidence interval of the study. A long line means an imprecise result, typically from a small study.
- 4. Weight. The percentage of the total weight the study receives, which depends on its precision, and under random effects, also on the between-study variance.
- 5. Diamond. The pooled estimate. Its center is the estimate and its width is the confidence interval.
- 6. Line of no effect. A vertical line at 1 for ratios and 0 for differences.
- 7. Direction labels. Text below the axis showing which side favors which intervention. Always check them, as they vary between figures.
Reading it, step by step
Read the axis and the labels
Check the scale and the direction. Here a risk ratio below 1 favors the intervention, but in another figure the labels may be reversed, so never assume.
Look at each study
Is the square left or right of the line of no effect, and how far? Do the intervals cross the line? In this example all six squares lie to the left of 1, and all six intervals cross it.
Compare the studies
Do the intervals overlap? Studies that agree have intervals that overlap broadly. Studies whose intervals barely overlap suggest heterogeneity.
Look at the weights
Which studies dominate? Here Study D carries 54 percent of the weight, and its estimate, 0.90, is the closest to no effect, so the pooled estimate sits toward it.
Read the diamond
The pooled risk ratio is 0.83, with a 95 percent interval from 0.72 to 0.95. The diamond does not cross 1, so the pooled result is statistically significant, although none of the six studies was significant on its own. That is the gain in precision from combining them.
Check the statistics
Look for heterogeneity statistics, the model used and the number of studies and participants, which are usually printed beneath the plot.
Why many plots use a log scale
For ratio measures the horizontal axis is almost always logarithmic, which is why the tick marks for 0.25, 0.5, 1 and 2 are equally spaced. On a logarithmic scale a halving and a doubling are the same distance from 1, which matches how ratios behave: a risk ratio of 0.5 and one of 2 are equally strong effects in opposite directions. On a plain scale, effects below 1 are squeezed into a small space and effects above 1 stretched out, which distorts the picture. The confidence intervals look symmetrical on the log scale for the same reason, and become asymmetrical when numbers are read off in the original units. For differences, such as mean differences, the axis is linear and symmetric about zero.
Where one row comes from
Each row of a forest plot is computed from the study's own data. Take Study A in the example, with a risk ratio of 0.72 and an interval from 0.45 to 1.15. On the log scale its estimate is -0.329, and because the interval is symmetric on that scale, its standard error can be recovered from the interval as the difference between the log of the upper and lower limits divided by 3.92, which gives 0.239. The inverse-variance weight is one divided by the square of that standard error, 17.5. Dividing by the total weight of all six studies, 213.4, gives the 8.2 percent shown on the plot. The same calculation applies to every row, and it shows why a study with a wide interval gets a small square: its standard error is large and its weight is the reciprocal of the square of that. This is also how a reviewer can recover the data needed for a meta-analysis from a published plot or table when the original numbers are not given. The data are simulated.
Reading plots for different effect measures
- Risk ratio, odds ratio, hazard ratio
- Ratios, with the line of no effect at 1 and a logarithmic axis. Values below 1 mean lower risk, odds or hazard in the intervention group, and values above 1 mean higher.
- Mean difference
- A difference in the original units, with the line of no effect at 0 and a linear axis. The sign depends on how the comparison is defined, and on whether a higher score is better, which differs by scale.
- Standardized mean difference
- A difference in standard deviation units, with no effect at 0. It lets studies that used different scales be combined, and its size is often described as small, medium or large by convention, which is a rough guide only.
- Correlation
- Plotted on the original scale or the Fisher z scale, with no effect at 0. The direction means a positive or negative association, not benefit or harm.
- Proportion
- A prevalence or event proportion, with no line of no effect, since there is no comparison. The pooled value is the estimate of interest, and the prediction interval shows the range between settings.
What to look for beyond the diamond
- Consistency. Do the studies point in the same direction and overlap? If some lie on opposite sides with intervals that do not overlap, the pooled estimate may be an average of conflicting results.
- Dominance. If one study carries most of the weight, the result is largely that study's result. Check whether it is at low risk of bias.
- Precision. Wide intervals, in the studies and the diamond, mean imprecise evidence, whatever the point estimate.
- Size of the effect. A significant result can be small. Ask whether the lower and upper ends of the diamond are practically important.
- Prediction interval. Some plots show a line for the prediction interval, which describes where the effect in a new study would fall and is usually wider than the diamond.
- Model. Fixed-effect and random-effects results look different, and the label should say which was used.
- Heterogeneity statistics. I squared and tau squared are usually printed. Large values warrant caution in interpreting the pooled result.
Variants you will meet
- Subgroup forest plots
- Studies are grouped by a characteristic, with a pooled estimate for each group and sometimes an overall estimate, and the test of difference between groups is printed.
- Cumulative forest plots
- Studies are ordered by date and the pooled estimate is recalculated as each is added, showing how the evidence accumulated.
- Leave-one-out plots
- The pooled estimate is shown with each study removed in turn, to reveal influential studies.
- Network forest plots
- Estimates for each treatment against a reference are shown, often with a ranking.
- Prediction-interval plots
- A line or a lighter band shows the prediction interval below or around the diamond.
- Bubble plots
- Not forest plots, but used with meta-regression: each study is a circle sized by weight against a moderator.
Common misreadings
- Reading the wrong direction. Always check the labels under the axis.
- Treating a diamond that touches the line as a clear finding. If the diamond crosses or touches the line of no effect, the pooled result is not statistically significant, though the interval shows what effects are compatible with the data.
- Counting studies. Counting squares on each side of the line is vote counting, which ignores weight and precision.
- Ignoring the scale. A risk ratio of 0.9 with a narrow interval can be trivial or important, depending on baseline risk.
- Equating a narrow diamond with truth. A narrow interval shows precision. If the studies are biased, the diamond is a precise estimate of a biased quantity.
- Overlooking heterogeneity. A tidy diamond over conflicting studies hides disagreement that the squares reveal.
Making a good forest plot
Authors preparing a plot should keep it legible and honest. The studies are ordered meaningfully, by year, by weight or by subgroup, and the order is stated. The scale and the direction are labeled. The model and the heterogeneity statistics are printed. The numbers for each estimate and interval are shown, so that the plot can be checked without measuring. The figure should be large enough to read at the journal's page size, with a font that is clear, and the squares should be sized by weight. Software packages offer many options, and defaults can mislead, for example by showing weights from a fixed-effect model beside a random-effects pooled estimate, so the output should be checked against the analysis. The figure is part of the reporting required by PRISMA 2020.
Support
Forest plots and the other figures of a meta-analysis are produced as part of the meta-analysis service.
Get a quoteSend your data and target journal.
Frequently asked questions
What does the size of the square mean?
It is proportional to the weight the study carries in the analysis, which mainly reflects its precision. Larger squares are studies with more influence on the pooled result.
What does it mean if the diamond touches the vertical line?
The pooled result is not statistically significant at the 5 percent level. The interval still shows the range of effects compatible with the data, which may include important benefit or harm.
Why is the axis logarithmic?
For ratio measures, a log scale makes a halving and a doubling the same distance from 1, which matches how ratios behave, and makes confidence intervals symmetric.
Why does the diamond look different when I switch between fixed and random effects?
Because the weights and the width differ. Random effects gives more equal weights and a wider interval when studies vary.
What is a prediction interval line on a forest plot?
A range in which the true effect in a new study is expected to fall. It is usually wider than the diamond when studies vary.
Can I judge a meta-analysis from its forest plot alone?
It reveals a lot, including consistency, weights and precision, but not how the studies were found or appraised, so it should be read with the review's methods.
References
- Lewis S, Clarke M. Forest plots: trying to see the wood and the trees. BMJ. 2001;322(7300):1479-1480.
- Deeks JJ, Higgins JPT, Altman DG, editors. Chapter 10: Analysing data and undertaking meta-analyses. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
- Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
- Anzures-Cabrera J, Higgins JPT. Graphical displays for meta-analysis: an overview with suggestions for practice. Res Synth Methods. 2010;1(1):66-80.
- IntHout J, Ioannidis JPA, Rovers MM, Goeman JJ. Plea for routinely presenting prediction intervals in meta-analysis. BMJ Open. 2016;6(7):e010247.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71