What the test does
A funnel plot shows each study's effect estimate against its precision. When there is no bias and no real differences between studies, the points form a symmetric funnel around the pooled estimate. When small studies report systematically larger effects than large ones, the funnel is lopsided. Egger's test, published by Egger and colleagues in 1997, tests that lopsidedness formally.
The idea is to fit a straight line to the relationship between a study's standardized effect and its precision. If the funnel is symmetric, the line passes through the origin. If it is asymmetric, the line has a non-zero intercept. The test asks whether the intercept differs from zero, so the quantity of interest is the intercept and not the slope.
It is described as a test for small-study effects because that is all it can detect. It finds an association between study size and effect. Whether that association reflects publication bias, a real difference between small and large studies, or chance is a separate judgment.
How the regression works
For each study, calculate two quantities. The first is the standardized effect, which is the effect estimate divided by its standard error. The second is the precision, which is the reciprocal of the standard error. Then fit an ordinary least squares regression of the standardized effect on the precision:
standardized effect = intercept + slope x precision
The slope estimates the underlying effect, in a form that is similar to a weighted estimate. The intercept measures asymmetry. A large study has high precision, so its position is dominated by the slope term. A small study has low precision, and its standardized effect is dominated by the intercept. If small studies systematically report larger effects than the line implies, the intercept is pulled away from zero.
The intercept has a standard error from the regression, and the test statistic is the intercept divided by that standard error, compared with a t distribution with k minus 2 degrees of freedom, where k is the number of studies. Many software packages offer this directly, and some use a variant that weights the regression, which is mathematically equivalent to regressing the effect on its standard error.
A worked example
The table shows 10 simulated studies, invented for illustration, with effect estimates on a continuous scale.
| Study | Standard error | Effect | Precision (1/SE) | Standardized effect |
|---|---|---|---|---|
| 1 | 0.08 | 0.22 | 12.50 | 2.75 |
| 2 | 0.11 | 0.30 | 9.09 | 2.73 |
| 3 | 0.14 | 0.18 | 7.14 | 1.29 |
| 4 | 0.17 | 0.28 | 5.88 | 1.65 |
| 5 | 0.20 | 0.38 | 5.00 | 1.90 |
| 6 | 0.24 | 0.22 | 4.17 | 0.92 |
| 7 | 0.28 | 0.45 | 3.57 | 1.61 |
| 8 | 0.33 | 0.35 | 3.03 | 1.06 |
| 9 | 0.38 | 0.60 | 2.63 | 1.58 |
| 10 | 0.45 | 0.50 | 2.22 | 1.11 |
Regressing the standardized effect on precision gives an intercept of 0.77 with a standard error of 0.26, a t statistic of 2.97 on 8 degrees of freedom. The two-sided 5 percent critical value for 8 degrees of freedom is 2.306, so the intercept differs significantly from zero.
The fixed-effect pooled estimate from all ten studies is 0.27. Restricting the calculation to the five most precise studies (standard error of 0.20 or less) gives 0.25. The smaller studies pull the overall estimate upward, which is the pattern the test has detected. The data are simulated, and the result does not say anything about a real body of evidence.
Interpreting a result
A significant intercept means that the relationship between effect size and precision is stronger than chance alone would readily produce. It does not say why. The causes of small-study effects include:
- Publication bias, where small studies with unimpressive results are less likely to be published.
- Selective outcome reporting, where small studies report the outcomes that happened to be favorable.
- Heterogeneity, where small studies were done in different populations or with more intensive interventions, so their true effects differ.
- Methodological quality, where small studies have higher risk of bias, which tends to inflate effects.
- Chance, especially with few studies.
- Artifacts of the effect measure, such as the correlation between an odds ratio and its standard error, discussed below.
A non-significant result is not evidence of no bias. The test has low power, particularly with fewer than about ten studies, and a symmetric funnel can still coexist with selective reporting. The result belongs in the discussion of the certainty of evidence, where it may support a decision to downgrade for publication bias, alongside the funnel plot, the searches that were done and what is known about unpublished trials.
Limits and common problems
Low power. With few studies, the test usually fails to detect real asymmetry. The Cochrane Handbook suggests that tests for funnel plot asymmetry should generally be used only when there are at least ten studies, and only when the studies are not all of similar size.
Odds ratios and other ratio measures. For binary outcomes analysed as odds ratios, the standard error of the log odds ratio is mathematically related to the odds ratio itself, particularly for large effects and for events that are common or rare. This built-in correlation can produce asymmetry where none exists, so the test may reject too often. Modified tests have been proposed to reduce the problem, and the choice should be checked in the software used.
Heterogeneity. When true effects differ between studies, the funnel is wider and the interpretation of asymmetry is harder. The test does not distinguish between asymmetry caused by bias and asymmetry caused by real differences.
Multiple testing. Running several asymmetry tests and reporting the one that supports a story is an analytic flexibility that should be avoided. Name the test in the protocol.
Alternatives to consider
Several other methods address the same question, and they behave differently.
- Begg and Mazumdar rank correlation test. Tests the correlation between standardized effect and variance. It has generally lower power than Egger's test.
- Harbord and Peters tests. Modifications for binary outcomes that reduce the artificial correlation described above.
- Contour-enhanced funnel plots. Shaded areas showing statistical significance help to judge whether missing studies are likely to be those with non-significant results.
- Trim-and-fill. Estimates the number of missing studies and adjusts the pooled estimate, as a sensitivity analysis.
- Selection models. Model the probability that a study is published as a function of its p value and adjust the estimate accordingly. They make assumptions that cannot be checked from the data alone.
- Limit meta-analysis and regression-based adjustments. Use the regression line to project the effect expected for an infinitely precise study.
No method corrects for publication bias reliably, because the missing studies are, by definition, not in the data. These are sensitivity analyses that show how much the conclusion depends on the assumption that nothing is missing.
How to report it
A clear report names the test and the number of studies, gives the intercept with its confidence interval or standard error and the p value, shows the funnel plot, and states the interpretation with its caveats. A sentence such as the following is typical: "Egger's test indicated funnel plot asymmetry (intercept 1.5, 95 percent CI ... , p = ...), which may reflect publication bias, heterogeneity or other small-study effects." Do not write that the test detected publication bias. Say that it detected asymmetry and discuss the causes.
Report what you did to look for unpublished or non-indexed studies, because a thorough search is a better protection against publication bias than any test applied afterwards. PRISMA 2020 asks authors to describe methods used to assess the risk of bias due to missing results in a synthesis and to present the results of that assessment.
Reading the intercept
The size of the intercept has a rough interpretation. An intercept of zero corresponds to a symmetric funnel. As the intercept moves away from zero, the funnel becomes more lopsided, and the sign shows the direction. In the example above the intercept is positive, which means the smaller, less precise studies report larger effects than the larger ones. Where a lower value of the outcome is the favorable direction, the same pattern would produce a negative intercept, so the sign must be read together with the coding of the outcome and not on its own.
The intercept is not an estimate of the amount of bias in the pooled effect. A bigger intercept does not translate into a specific correction. For that reason it is poor practice to report the intercept as a measure of how much the result is inflated. If a quantitative adjustment is wanted, it comes from a separate sensitivity analysis, with its own assumptions stated, and the unadjusted estimate remains the main result.
Planning the test in a protocol
Decisions about asymmetry testing are best made before the data are seen. A protocol can state the minimum number of studies for which a test will be run, the test to be used for each effect measure, the significance level, and what will be done if the test is positive. Registering these choices guards against the temptation to try several tests and to report the one that gives the preferred answer. It also settles in advance the practical question of what to do with a review of fewer than ten studies. The honest answer is usually that no test will be run, that the funnel plot will be shown only if it is informative, and that the possibility of missing studies will be addressed by the search and in the discussion.
It helps to state how the result will feed into the assessment of the certainty of evidence. In GRADE, concern about publication bias is one of the domains that can lower certainty, and a transparent rule, such as considering a downgrade when asymmetry is present and the search had limitations, is easier to defend than a judgment made after the fact.
How we can help
We can run asymmetry tests and contour-enhanced funnel plots, choose a test that suits the effect measure, carry out sensitivity analyses such as trim-and-fill and selection models, and write the limitations clearly. [OWNER VERIFICATION REQUIRED] The related services are meta-analysis, statistical analysis and risk of bias and certainty of evidence.
Frequently asked questions
What does Egger's test measure?
Whether the relationship between study size and effect is stronger than chance would explain, which shows up as funnel plot asymmetry. It tests whether the regression intercept differs from zero.
How many studies are needed?
About ten or more, and not all of similar size. With fewer studies the test has low power, and a non-significant result is uninformative.
Does a significant result prove publication bias?
No. It indicates small-study effects, which can also arise from heterogeneity, differences in study quality, selective reporting or chance.
Is Egger's test suitable for odds ratios?
It can give false positives, because the standard error of a log odds ratio is related to the odds ratio. Modified tests, such as those by Harbord or Peters, are preferred for binary outcomes.
What if the test is not significant?
That does not show the absence of bias, especially with few studies. Consider the funnel plot, the search and other evidence about unpublished studies.
What should I do after a significant result?
Examine the funnel plot, consider causes other than publication bias, run sensitivity analyses such as trim-and-fill, and consider downgrading the certainty of evidence.
References
- Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315(7109):629-634.
- Sterne JAC, Sutton AJ, Ioannidis JPA, et al. Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ. 2011;343:d4002.
- Begg CB, Mazumdar M. Operating characteristics of a rank correlation test for publication bias. Biometrics. 1994;50(4):1088-1101.
- Harbord RM, Egger M, Sterne JAC. A modified test for small-study effects in meta-analyses of controlled trials with binary endpoints. Stat Med. 2006;25(20):3443-3457.
- Peters JL, Sutton AJ, Jones DR, Abrams KR, Rushton L. Comparison of two methods to detect publication bias in meta-analysis. JAMA. 2006;295(6):676-680.
- Page MJ, Higgins JPT, Sterne JAC. Chapter 13: Assessing risk of bias due to missing results in a synthesis. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71