What the method is for
A diagnostic test is used to decide whether a person has a condition. How well it does so is described by its accuracy, estimated in studies that apply the test, called the index test, and a reference standard, which is the best available way of determining the true condition, to the same people. A systematic review of test accuracy collects such studies, and a meta-analysis combines their results to estimate how accurate the test is on average, how much accuracy varies, and why. The same approach is used for screening tests, imaging, laboratory assays, clinical signs and decision rules.
The method resembles ordinary meta-analysis in its aims and differs in its statistics, because a study of test accuracy yields two linked results, sensitivity and specificity, and not one effect. It also differs in its appraisal, since the biases that matter, such as the way patients were selected and the way the reference standard was applied, are specific to accuracy studies. The general principles are in the page on meta-analysis, and the practice is covered in the meta-analysis service.
Accuracy measures
Each study provides a two-by-two table of the test result against the reference standard: true positives, false negatives, false positives and true negatives. From it come the main measures. Sensitivity is the proportion of people with the condition whom the test identifies as positive. Specificity is the proportion of people without the condition whom the test identifies as negative. Positive and negative likelihood ratios express how much a positive or a negative result changes the odds of disease, and the diagnostic odds ratio summarizes accuracy in one number but conceals the balance between sensitivity and specificity. Predictive values depend on the prevalence of the condition in the people tested, so they cannot be carried over to other settings without recalculation.
A simple example shows the arithmetic. Suppose a test is applied to 300 people, 100 of whom have the condition. It correctly identifies 90 of them (true positives) and misses 10 (false negatives), and among the 200 without the condition it correctly clears 170 (true negatives) and wrongly flags 30 (false positives). Sensitivity is 90/100 = 0.90, specificity is 170/200 = 0.85, the positive likelihood ratio is 0.90/0.15 = 6.0, the negative likelihood ratio is 0.10/0.85 = 0.118, and the diagnostic odds ratio is (90 x 170)/(10 x 30) = 51. Prevalence in this sample is 0.33, so the positive predictive value is 0.75 and the negative predictive value is 0.94, and both would change if the test were used in a population with a different prevalence. The figures are invented to illustrate the calculation.
The threshold effect
Many tests give a continuous result, such as a biomarker level or a score, and the test is called positive above a threshold. Different studies may use different thresholds, which affects accuracy in a predictable way: a lower threshold raises sensitivity and lowers specificity, and a higher one does the opposite. Sensitivity and specificity therefore tend to be negatively correlated across studies, and that correlation reflects the trade-off along the receiver operating characteristic curve. Pooling sensitivity and specificity separately, as if they were independent, ignores this and can give a misleading summary that does not correspond to any actual operating point.
The correct statistical response is to model the two quantities together. The analyst also records the threshold used in each study, examines whether a threshold effect is present, for example from a plot of sensitivity against specificity in receiver operating characteristic space, and explores it with threshold as a covariate. Simple correlation of the logit transformed measures is sometimes used to detect it, though more informative approaches are model-based.
Bivariate and hierarchical models
Two closely related models are the standard. The bivariate random-effects model works on the logit of sensitivity and the logit of specificity in each study, which are modeled as drawn from a bivariate normal distribution whose means are the summary logits and whose covariance includes the between-study variance of each and their correlation. It allows for heterogeneity and the threshold effect and can be fitted as a generalized linear mixed model with exact binomial likelihoods, which handles studies with extreme values without continuity corrections. The hierarchical summary receiver operating characteristic model parametrizes accuracy and threshold separately, giving a summary curve, and the two are mathematically equivalent in the absence of covariates and can be translated into each other.
The bivariate model is preferred when the interest is in a summary sensitivity and specificity at a common threshold, and the hierarchical model when the aim is a summary curve across thresholds. Either can include covariates, such as study design or test version, to explore heterogeneity. Results are reported as pooled sensitivity and specificity with confidence regions, and a summary curve with a prediction region that shows the range expected in a new study. Univariate pooling and simple methods that ignore the correlation are not recommended, and a Moses-Littenberg approach, once common, is superseded.
Heterogeneity
Heterogeneity in accuracy studies is the rule. Beyond the threshold effect, accuracy varies with the spectrum of patients, which is the mix of disease severity and of the other conditions that can be mistaken for the target one, the setting, the version and the operator of the test, and the reference standard. Studies done in referral centers with patients who clearly have or clearly lack the condition tend to give higher estimates than those in primary care, a phenomenon known as spectrum bias. For these reasons, I squared as used in intervention reviews is not appropriate for accuracy data, and heterogeneity is assessed by visual inspection of the forest plots of sensitivity and specificity, the dispersion in receiver operating characteristic space, and the between-study variances in the model.
Sources of heterogeneity are explored with covariates or subgroup analyses, which need enough studies, as in any meta-regression. See meta-regression.
Risk of bias and applicability
QUADAS-2 is the standard tool for assessing the quality of accuracy studies. It has four domains. Patient selection asks whether the patients were enrolled consecutively or randomly and whether inappropriate exclusions or a case-control design were used, since the latter inflates accuracy. The index test domain asks whether the test was interpreted without knowledge of the reference standard and whether any threshold was prespecified. The reference standard domain asks whether it is likely to classify the condition correctly and whether it was interpreted blind to the index test. Flow and timing asks whether the interval between tests was appropriate, whether all patients received the same reference standard and whether all were included in the analysis. Each domain is rated for risk of bias, and the first three are also rated for concerns about applicability to the review question. See risk-of-bias tools.
A reference standard is never perfect, and an imperfect standard biases estimates of accuracy in a direction that depends on its errors. Methods for adjusting exist but require assumptions, and the problem is usually discussed as a limitation. Verification bias occurs when the reference standard is applied selectively according to the index test result.
Reporting bias
Publication bias is harder to assess in accuracy reviews. The funnel plots and tests used for treatment effects can be misleading because the sample size is related to the diagnostic odds ratio in ways that are not due to bias. A modified approach, the Deeks funnel plot asymmetry test, uses the effective sample size and is recommended when it is assessed at all, but its power is low, and it is interpreted cautiously. The search should include registries and grey literature where appropriate, and the discussion should consider whether unpublished studies with poor accuracy might be missing.
Conducting a review
Frame the question
The population, the index test with its threshold, the target condition and the reference standard are specified, along with the role of the test in the care pathway, since accuracy relevant to triage differs from that relevant to confirmation.
Search and select
Searches avoid study-design filters, which perform poorly for accuracy studies, and use concepts for the test and the condition. Selection is done in duplicate.
Extract and appraise
Two-by-two data are extracted for each threshold reported, with study and patient characteristics, and QUADAS-2 is applied.
Analyze
Forest plots, a plot in receiver operating characteristic space, the bivariate or hierarchical model, and covariate analyses are run.
Assess certainty and report
The certainty of evidence is rated with an approach adapted for accuracy, and the review is reported with PRISMA-DTA.
Presenting and interpreting results
Results are presented with paired forest plots of sensitivity and specificity, a summary point with its confidence region in receiver operating characteristic space, and, where thresholds vary, a summary curve. The most useful interpretation translates them into consequences for a specified population: for a stated prevalence, how many of 1,000 people tested would be correctly identified, how many missed and how many falsely alarmed. Such natural-frequency summaries are far easier for clinicians and patients to understand than percentages, and they show why the same test performs differently in populations with different prevalence. Accuracy is not the same as clinical usefulness, which also depends on the consequences of the test results for patient outcomes, a matter that accuracy studies alone do not address.
Reporting
The review is reported with PRISMA 2020 and its extension for diagnostic test accuracy, PRISMA-DTA, which adds items about the index test, the reference standard, the target condition and the accuracy measures. The studies themselves should have followed STARD, and review authors can use STARD compliance as a marker of reporting quality. The report presents the data for each study so that the analysis can be repeated. See PRISMA extensions.
Limitations
The method is only as good as the primary studies, and many accuracy studies are at high risk of bias, especially for patient selection. Heterogeneity is large and often unexplained. Few studies may use each threshold, so thresholds are pooled, which blurs results. An imperfect reference standard biases estimates. The models are more complex than standard meta-analysis, with convergence issues when studies are few, and sparse data are common. And accuracy does not equal benefit: a test with high accuracy may not improve outcomes. The results are not advice on the care of individuals.
Diagnostic accuracy reviews support evaluation of tests. They do not provide clinical advice.
How we can help
Support for diagnostic accuracy meta-analysis
The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.
Feasibility check
A review of your studies and data to confirm that the method is suitable and which approach fits.
Analysis and figures
The analysis run to a prespecified plan, with forest plots and the other figures.
Methods and results text
Written for the manuscript and aligned with PRISMA 2020 or the relevant extension.
Manuscript and submission
Optional: the full paper, the reporting checklist and the submission materials.
Frequently asked questions
Why can't I pool sensitivity and specificity separately?
Because they are linked through the test threshold. Studies that use a lower threshold have higher sensitivity and lower specificity, so the two are correlated across studies, and separate pooling ignores that and can give a summary that corresponds to no real operating point.
What is the difference between the bivariate and the HSROC model?
They are equivalent in the absence of covariates. The bivariate model gives a summary sensitivity and specificity at a common threshold, and the HSROC model parametrizes accuracy and threshold to give a summary curve.
Why not use I-squared for accuracy data?
It was designed for a single effect measure and does not account for the correlation between sensitivity and specificity or the threshold effect. Heterogeneity is assessed with plots and the variances in the bivariate model.
What is spectrum bias?
Variation in accuracy with the mix of patients, for example between a referral clinic and primary care, and the tendency of case-control designs to overestimate accuracy.
Are predictive values suitable for pooling?
They depend on prevalence, so they should not be pooled across settings. They can be computed for a stated prevalence from pooled sensitivity and specificity.
Which reporting guideline applies?
PRISMA-DTA for reviews of diagnostic accuracy, with the primary studies expected to follow STARD.
References
- McInnes MDF, Moher D, Thombs BD, et al. Preferred reporting items for a systematic review and meta-analysis of diagnostic test accuracy studies: the PRISMA-DTA statement. JAMA. 2018;319(4):388-396.
- Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
- Bossuyt PM, Reitsma JB, Bruns DE, et al. STARD 2015: an updated list of essential items for reporting diagnostic accuracy studies. BMJ. 2015;351:h5527.
- Reitsma JB, Glas AS, Rutjes AWS, Scholten RJPM, Bossuyt PM, Zwinderman AH. Bivariate analysis of sensitivity and specificity produces informative summary measures in diagnostic reviews. J Clin Epidemiol. 2005;58(10):982-990.
- Rutter CM, Gatsonis CA. A hierarchical regression approach to meta-analysis of diagnostic test accuracy evaluations. Stat Med. 2001;20(19):2865-2884.
- Harbord RM, Deeks JJ, Egger M, Whiting P, Sterne JAC. A unification of models for meta-analysis of diagnostic accuracy studies. Biostatistics. 2007;8(2):239-251.
- Chu H, Cole SR. Bivariate meta-analysis of sensitivity and specificity with sparse data: a generalized linear mixed model approach. J Clin Epidemiol. 2006;59(12):1331-1332.
- Deeks JJ. Systematic reviews of evaluations of diagnostic and screening tests. BMJ. 2001;323(7305):157-162.
- Deeks JJ, Macaskill P, Irwig L. The performance of tests of publication bias and other sample size effects in systematic reviews of diagnostic test accuracy was assessed. J Clin Epidemiol. 2005;58(9):882-893.
- Leeflang MMG, Rutjes AWS, Reitsma JB, Hooft L, Bossuyt PM. Variation of a test's sensitivity and specificity with disease prevalence. CMAJ. 2013;185(11):E537-E544.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71