Four kinds of missing data
The phrase covers different problems, and treating them as one leads to the wrong remedy.
- Missing studies. Eligible studies were never found or never published. This is the problem of publication bias and incomplete searching, handled by searching widely and examining small-study effects.
- Missing outcomes. A study collected an outcome but did not report it, often because the result was not significant. This is selective outcome reporting, and it biases the studies that remain.
- Missing summary statistics. A study reports the outcome but not the numbers needed, such as a standard deviation or the number of participants analysed. These can sometimes be derived or obtained from the authors.
- Missing participants. Within a study, some participants were lost to follow-up or excluded, so outcomes are unknown for them. This affects the validity of each study's result.
Missing studies and missing outcomes
For missing studies, the protection is in the search: databases, registries, regulatory documents and contact with investigators. After the search, funnel plots and asymmetry tests, with their limits, give some indication, and sensitivity analyses such as trim-and-fill and selection models show how much a conclusion depends on the assumption that nothing is missing.
For missing outcomes, compare each study's report with its registry entry or protocol. When an outcome was planned but not reported, contact the authors. The Outcome Reporting Bias in Trials (ORBIT) approach classifies the likelihood that an outcome was withheld, and sensitivity analyses can test the effect of plausible missing results. The risk-of-bias tool for randomized trials, RoB 2, has a domain on selection of the reported result, and the Cochrane guidance on missing results in a synthesis, in Chapter 13 of the Handbook, covers both problems. Where many studies in a meta-analysis lack the outcome, say so, because the pooled estimate then rests on a selected subset.
Missing summary statistics
Standard deviations are most often missing. The preferred remedy is to calculate them from other reported values: standard errors, confidence intervals, t statistics or p values, ranges or interquartile ranges, as described in the guide to data extraction. If these are not available, contact the authors. As a last resort, a standard deviation can be borrowed from another study in the review that used the same measure, or imputed from the average of the others, with a sensitivity analysis to show that the result does not depend on the borrowed value. Such studies should be marked in the forest plot or table, and should never be silently treated as if they reported the value.
For change-from-baseline outcomes, the standard deviation of the change needs the baseline-follow-up correlation, which is almost never reported. Take it from a study in the review that does report it, assume a plausible value such as 0.5 and examine the effect of other values.
Missing sample sizes are handled by using the number randomized, with a note. An effect estimate without a measure of precision cannot be weighted, and a study that supplies only a direction and a statement of significance cannot be pooled, although it can be described.
Missing participants within a study
When participants drop out, the outcome is unknown for them, and an analysis that ignores them can be biased, especially if dropout differs between arms or is related to the outcome. Trials report this in different ways: complete-case analysis, last observation carried forward, multiple imputation, or an intention-to-treat analysis with some assumption. In the review, the first step is to record the numbers randomized, analysed and lost in each arm, and to judge risk of bias from attrition. The second is to ask whether the trial's chosen approach is reasonable.
A sensitivity analysis can show how much conclusions depend on assumptions about the missing participants. Take a simulated trial with 100 randomized to each arm. In the treatment arm 10 participants (10 percent) are missing, and 20 of the 90 observed have the event. In the control arm 15 (15 percent) are missing, and 30 of the 85 observed have the event.
- Complete case. Risk ratio = (20/90) / (30/85) = 0.63.
- Favorable to treatment. Assume no event in the missing treated participants and an event in all missing controls: (20/100) / (45/100) = 0.44.
- Unfavorable to treatment. Assume an event in all missing treated and none in the missing controls: (30/100) / (30/100) = 1.00.
The extreme cases bracket the plausible range, from 0.44 to 1.00. The complete-case result lies between them. If the conclusion is the same under both extremes, missing participants are unlikely to alter it. If the extremes straddle no effect, as they do here, the result is sensitive to missing data. These extreme assumptions are unrealistic, since not all missing participants would have had the event, and intermediate approaches exist. One uses the idea of an informative missingness odds ratio, which expresses how different the odds of an event are in missing participants compared with observed ones, and allows a range of plausible values. The data are simulated. The same logic applies to any quantity that must be assumed: state the assumption, show the range tested and report how far the result moved. Readers can then judge for themselves whether the conclusion is robust, and a reviewer can reproduce the analysis from the stated values.
Imputation and its cautions
Multiple imputation, in which several plausible values are generated for each missing value and the results are combined, is a standard approach in primary studies when data are missing at random, meaning that missingness depends only on observed information. It is less often possible at the review level because the individual participant data are not available. A review mostly has to rely on what the trial authors did, and on sensitivity analyses of the sort above. If individual participant data are obtained, imputation can be done on the combined dataset.
Two practices should be avoided. The first is last observation carried forward as a primary analysis, which assumes that participants who dropped out stayed unchanged, and which can bias results in either direction. The second is replacing missing outcomes with a favorable or unfavorable value without saying so. Any assumption should be explicit.
What to report
State, for each type of missing data, how many studies were affected, what attempts were made to obtain data, what was imputed or derived and how, and the results of sensitivity analyses. Report attrition per arm in the characteristics table, and indicate in the forest plot which results used derived or imputed values. Link the findings to the risk-of-bias judgments and to the certainty ratings: substantial attrition, selective reporting and missing studies are each reasons to consider lowering certainty in GRADE. A frank statement of the limitation is better than a sophisticated method that conceals it. A short table listing each study with missing information, the type of missing data, the action taken and the result of the sensitivity analysis gives readers a complete picture in one place, and it also makes the review easier to update when authors supply data later. If new data arrive after the analysis, rerun the sensitivity analyses and update the table, noting the date of the change.
An example of borrowing a standard deviation
Four studies of the same scale report standard deviations of 8.1, 7.4, 9.0 and 8.6. A fifth, with 60 participants per arm, reports a mean difference but no standard deviation. Imputing the average, 8.28, gives a standard error for the difference of 8.28 x sqrt(2/60) = 1.51. Using the smallest and largest observed values gives 1.35 and 1.64, so the study's weight, which depends on the inverse of the squared standard error, would change from about 15 percent lower to about 25 percent higher than with the average value. If the pooled result is stable across that range, the imputation is unlikely to matter. If it is not, the study should not be allowed to drive the conclusion. The values are simulated.
Contacting authors effectively
A good request is short, specific and easy to answer. Name the review and its registration, list exactly which numbers are needed, for which outcome, group and time point, say how they will be used and offer an alternative format such as a table that can be filled in. Send it to the corresponding author, and to the last author if there is no reply. Allow a reasonable time, send a reminder once, and keep a log of requests and replies. Response rates vary widely and many requests go unanswered, which should be reported rather than hidden. Where a reply provides data that differ from the published paper, record both and say which was used. Be aware of data-sharing agreements and of the restrictions of confidentiality when obtaining participant-level information.
Choosing a response by situation
- A study reports the outcome without a standard deviation. Derive it, ask the authors, then borrow with a sensitivity analysis.
- A study clearly measured an outcome but did not report it. Check the registry, ask the authors, classify the risk of selective reporting and consider the effect on the pooled result.
- A study has high dropout. Judge risk of bias, examine the trial's own handling of missing data and run extreme-case sensitivity analyses.
- Many studies lack the outcome. Present the available data, say that the pooled estimate is based on a subset, and be cautious about conclusions.
- A study of the right design is not found. Search registries and contact experts, and assess small-study effects.
How we can help
We can contact authors, derive missing statistics with documented methods, classify the risk of missing outcomes, run sensitivity analyses for missing participants and report the effect on conclusions and certainty. [OWNER VERIFICATION REQUIRED] The relevant services are meta-analysis, data extraction and coding and risk of bias and certainty of evidence.
Frequently asked questions
What types of missing data occur in a review?
Missing studies, missing outcomes, missing summary statistics and missing participants. Each needs a different response.
What can I do if a standard deviation is missing?
Calculate it from other reported statistics, ask the authors, or borrow a value from similar studies as a last resort, with a sensitivity analysis.
How do I check for selective outcome reporting?
Compare the published report with the registry entry or protocol, contact the authors about unreported outcomes and consider an approach such as ORBIT.
What is a best-case worst-case analysis?
A sensitivity analysis that assigns the most favorable or unfavorable outcome to missing participants, to see whether the conclusion depends on them.
Should I use last observation carried forward?
Not as a primary approach, because it assumes dropouts stay unchanged and can bias results.
How does missing data affect certainty of evidence?
Substantial attrition and selective reporting can lower certainty through risk of bias, and missing studies can lower it through publication bias.
References
- Higgins JPT, Li T, Deeks JJ. Chapter 6: Choosing effect measures and computing estimates of effect. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Page MJ, Higgins JPT, Sterne JAC. Chapter 13: Assessing risk of bias due to missing results in a synthesis. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Kirkham JJ, Dwan KM, Altman DG, et al. The impact of outcome reporting bias in randomised controlled trials on a cohort of systematic reviews. BMJ. 2010;340:c365.
- Dwan K, Gamble C, Williamson PR, Kirkham JJ. Systematic review of the empirical evidence of study publication bias and outcome reporting bias: an updated review. PLoS One. 2013;8(7):e66844.
- White IR, Higgins JPT, Wood AM. Allowing for uncertainty due to missing data in meta-analysis. Part 1: two-stage methods. Stat Med. 2008;27(5):711-727.
- Akl EA, Briel M, You JJ, et al. Potential impact on estimated treatment effects of information lost to follow-up in randomised controlled trials (LOST-IT): systematic review. BMJ. 2012;344:e2809.
- Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.