What publication bias is
Not every study that is carried out is published, and the ones that are not are not a random subset. Studies with statistically significant results, large effects or results in the expected direction are more likely to be written up, submitted, accepted and published, and to be published sooner and in more prominent places. Studies that find nothing, or that find the opposite of what was hoped, are more often left in a drawer. A review that draws only on the published literature will then see a distorted picture: effects look larger and more consistent than they really are.
The problem was recognized long ago and has been documented repeatedly. Comparisons of trials registered with regulators or ethics committees with those later published have found that trials with significant results were more likely to appear in print, and that when outcomes were reported selectively, the ones reported were the ones that favored the intervention. It affects fields beyond medicine, from psychology to economics. It is one of the main reasons for the care taken over searching in a systematic review, described in systematic review.
The forms of bias in reporting
- Publication bias
- Whole studies are not published, depending on their results.
- Time-lag bias
- Studies with positive results are published sooner, so early reviews overstate effects.
- Selective outcome reporting
- Within a published study, only some of the measured outcomes or analyses are reported, favoring those with significant results.
- Language and location bias
- Studies in some languages, or published in less accessible journals, are less likely to be found, and may differ in results from those in the databases searched.
- Citation and duplicate-publication bias
- Studies with positive results are cited more, and are more likely to be reported several times, which affects what is found and counted.
Together they are called reporting biases, and their effect on a meta-analysis is similar: a distortion that standard analysis cannot detect from the studies alone. Because selective outcome reporting occurs inside the studies that were found, it is assessed with the risk-of-bias tools and by comparison with protocols and registries. See risk-of-bias tools.
Funnel plots
The funnel plot is the standard graphical check. It plots each study's effect estimate on the horizontal axis against a measure of its precision, usually the standard error, on the vertical axis, with the most precise studies at the top. In the absence of bias and heterogeneity, the studies scatter symmetrically around the pooled estimate, in an inverted funnel: large, precise studies cluster near the top close to the true effect, and small, imprecise ones spread widely at the bottom, with random errors in both directions. If small studies with unfavorable results are missing, a corner of the funnel is empty and the plot looks lopsided.
Interpretation is subjective, and visual assessment is unreliable. Asymmetry has causes other than publication bias, so it is called a small-study effect and treated as a signal and not as proof. Differences in the populations or interventions of small and large studies, poor methods in small studies, a choice of effect measure that is related to its standard error, and chance can all produce it. For these reasons funnel plots are recommended only when there are about ten or more studies, and they should be read with those alternatives in mind. See funnel plots.
Egger's test: a worked example
Egger's regression test turns the funnel plot into a calculation. Each study's standardized effect, the estimate divided by its standard error, is regressed on its precision, the reciprocal of the standard error. If there is no asymmetry, the line passes through the origin, so that the intercept is zero. A positive or negative intercept that differs from zero beyond chance suggests asymmetry, with larger magnitude meaning more asymmetry.
Here is a simulated example, invented for illustration, of 10 studies in which smaller studies, with larger standard errors, tend to show larger effects. The standard errors run from 0.08 to 0.35 and the effects from 0.18 to 0.66. The regression gives an intercept of 1.39 with a standard error of 0.31, so that t is 4.54 on 8 degrees of freedom, well beyond the critical value of 2.31 at the 5 percent level. The precision-weighted pooled estimate across all studies is 0.28, while the estimate from the 5 most precise studies alone is 0.25, which is smaller. The pattern is typical: the pooled result is pulled upward by small studies with large effects. The result does not tell us why, and a real analysis would ask whether small studies differ in population or method, or whether small negative studies might be missing.
The test has weaknesses. It has low power with fewer than about ten studies, so a non-significant result is not reassuring. For some effect measures, notably the log odds ratio, the standard error is related to the effect estimate even without bias, which can give false positives, and alternative tests designed for binary outcomes have been proposed. See Egger's test.
Methods that adjust for it
Several methods try to adjust the pooled estimate for suspected bias. The trim-and-fill method estimates the number of studies missing from one side of the funnel, adds imputed mirror-image studies, and recalculates the pooled effect. Selection models describe the probability that a study is published as a function of its p-value or its size and estimate the effect corrected for it. Regression-based approaches extrapolate the relationship between effect and precision to a study with no sampling error, and give an estimate of the effect that a very large study would find. Each rests on assumptions that cannot be verified from the data, and in simulations they perform unevenly, especially with heterogeneity.
For this reason they are best used as sensitivity analyses, to see how much the conclusion might change under a particular pattern of selection, and not as corrections to be trusted. A pooled estimate that remains clinically important across several such analyses is reassuring, and one that vanishes under them is a warning. See trim-and-fill.
Prevention is better than correction
Since no statistical method reliably repairs publication bias, the main defense is to avoid it at the source. A systematic review can search for unpublished material, including trial registries, regulatory documents, conference abstracts, theses and preprint servers, and it can contact investigators. It can compare registered outcomes with published ones to detect selective reporting. Registration of trials, and of reviews, before the work starts, makes the existence of studies visible whatever their outcome. Journals that evaluate studies on the quality of their methods and not the direction of their results, and registered reports, in which a study is accepted for publication on the basis of its protocol, reduce the pressure that creates the problem. And a culture that values reporting null results helps. The search methods are described in grey literature.
Publication bias in the assessment of certainty
In GRADE, a suspicion of publication bias is one of the reasons to lower the certainty of evidence for an outcome. The guidance recognizes that it is hard to prove, and suggests considering the pattern of the evidence: whether it consists of a few small studies with positive results, particularly from commercial sponsors; whether funnel plots and tests suggest asymmetry; whether searches of registries have found unpublished studies; and whether there is a reason to expect results to be withheld. A rating is lowered when the likelihood is judged to be high enough to matter. Where the evidence is large and consistent, and searches for unpublished studies found little, the concern is less. See GRADE.
Selective outcome reporting in practice
Selective outcome reporting is often more damaging than the loss of whole studies, because it happens inside studies that appear in the review. A trial may measure a dozen outcomes and publish the three that were significant, or switch the primary outcome after the data are in, or report an analysis that was not planned. Evidence from comparisons of protocols with publications has found that a substantial proportion of trials have outcome discrepancies, and that significant outcomes are more likely to be fully reported than non-significant ones. The check is to compare the outcomes in the registry entry or protocol with those in the paper, which is possible when the trial was registered. When an expected outcome is missing, the reviewer asks whether it was measured, and, if the answer is unclear, contacts the authors. If an outcome that should have been measured is missing from many trials, this pattern is itself a warning that results may have been withheld. The risk-of-bias tool for randomized trials has a domain on selection of the reported result for this reason.
Sponsorship, regulatory data and unpublished results
Industry-sponsored studies have been found in several comparisons to be more likely to report favorable results, and sponsors may withhold unfavorable ones. Regulatory agencies hold detailed reports of the studies submitted for drug approval, including those that were never published, and reviews that have used these documents have sometimes reached different conclusions from reviews of the published literature. Access to them has improved in some jurisdictions. Results posted in trial registries, which in some countries are required by law within a set time of completion, provide another route to unpublished results. A review of a pharmacological or device question should therefore search for regulatory and registry results, and say what was found. Where such data are not available, the possibility that unfavorable results are missing should be considered in the interpretation, and it contributes to the judgment about publication bias in the certainty of evidence.
Reporting
A review should describe how it sought unpublished studies and how it assessed reporting bias: which sources were searched beyond the main databases, whether registries were checked, how many studies were included, and which methods of assessment were used. It should present a funnel plot and the test, if there are enough studies, and any sensitivity analyses, with their assumptions, and it should state its conclusion about the possible effect of bias on the result. Honest reporting acknowledges that the possibility cannot be excluded. PRISMA 2020 includes items for the methods and results of assessing reporting bias. See PRISMA 2020.
Common mistakes
- Testing for asymmetry with fewer than ten studies and treating a non-significant test as proof of no bias.
- Calling any asymmetry publication bias, ignoring other explanations.
- Reporting an adjusted estimate from trim-and-fill as the corrected effect.
- Relying on the visual impression of a funnel plot without a test or a second opinion.
- Assessing only publication bias and ignoring selective outcome reporting within studies.
- Searching only the main databases and concluding that no unpublished studies exist.
Support
Assessment of reporting bias, searching for unpublished studies and the sensitivity analyses described here are part of the meta-analysis and search strategy services.
Get a quoteTell us your topic and the studies you have.
Frequently asked questions
What is the difference between publication bias and small-study effects?
Small-study effects are the pattern in which smaller studies show different, often larger, effects than larger ones. Publication bias is one possible cause. Others include real differences between small and large studies and poorer methods in small ones.
How many studies do I need to test for publication bias?
About ten or more is the usual recommendation. With fewer, tests have low power and funnel plots are hard to read.
Does a symmetrical funnel plot rule out publication bias?
No. It is reassuring but not conclusive, particularly with few studies, and selective outcome reporting is not visible in it.
Should I report the trim-and-fill estimate?
As a sensitivity analysis, with its assumptions, but not as the corrected effect.
How do I search for unpublished studies?
Search trial registries, regulatory documents, conference abstracts, theses and preprints, and contact investigators where appropriate.
Can I fix publication bias statistically?
Not reliably. Statistical methods show how sensitive a result is to assumed patterns of selection. Prevention and comprehensive searching are the better remedies.
References
- Page MJ, Higgins JPT, Sterne JAC. Chapter 13: Assessing risk of bias due to missing results in a synthesis. In: Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
- Sterne JAC, Sutton AJ, Ioannidis JPA, et al. Recommendations for examining and interpreting funnel plot asymmetry in meta-analyses of randomised controlled trials. BMJ. 2011;343:d4002.
- Egger M, Davey Smith G, Schneider M, Minder C. Bias in meta-analysis detected by a simple, graphical test. BMJ. 1997;315(7109):629-634.
- Duval S, Tweedie R. Trim and fill: a simple funnel-plot-based method of testing and adjusting for publication bias in meta-analysis. Biometrics. 2000;56(2):455-463.
- Dickersin K. The existence of publication bias and risk factors for its occurrence. JAMA. 1990;263(10):1385-1389.
- Dwan K, Gamble C, Williamson PR, Kirkham JJ; Reporting Bias Group. Systematic review of the empirical evidence of study publication bias and outcome reporting bias: an updated review. PLoS One. 2013;8(7):e66844.
- Hopewell S, Loudon K, Clarke MJ, Oxman AD, Dickersin K. Publication bias in clinical trials due to statistical significance or direction of trial results. Cochrane Database Syst Rev. 2009;(1):MR000006.
- Peters JL, Sutton AJ, Jones DR, Abrams KR, Rushton L. Comparison of two methods to detect publication bias in meta-analysis. JAMA. 2006;295(6):676-680.
- Guyatt GH, Oxman AD, Montori V, et al. GRADE guidelines: 5. Rating the quality of evidence: publication bias. J Clin Epidemiol. 2011;64(12):1277-1282.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71