Deduplicating the records
Searches in several databases return the same studies many times, in slightly different forms. Duplicates must be removed before screening, because screening the same record twice wastes effort and risks inconsistent decisions. Two different problems arise.
Exact and near duplicates across databases. The same article appears in MEDLINE, Embase and Scopus with small differences in the title, author names or page numbers. Reference managers such as EndNote, Zotero and Mendeley, and review tools such as Covidence, Rayyan, DistillerSR and EPPI-Reviewer, have deduplication functions. They match on fields such as title, year, journal, volume and pages, or DOI. Automatic matching misses some duplicates and, if set too loosely, removes records that are different, such as two reports from the same author group in the same year with similar titles.
Multiple reports of the same study. A trial can be described in a protocol, a main results paper, secondary analyses, a conference abstract and a registry entry. These are not duplicates in the technical sense, but they belong to one study and should be linked before analysis, otherwise the same participants are counted more than once. The unit of analysis in the review is the study, not the report.
Good practice is to run automatic deduplication, review the borderline matches manually, and keep a record of the number removed at each step. Some teams keep the removed records in a separate folder in case of dispute. Number the sources of each record before merging, since PRISMA 2020 asks for the number of records identified from each source.
Eligibility criteria before screening
Screening applies eligibility criteria, and the criteria must be written down before anyone screens. They usually follow the PICO structure, together with design, time period, language and publication status. Vague criteria cause inconsistency and later disagreement. A useful criterion is one that two people can apply and reach the same answer. For example, "adults with type 2 diabetes" needs a decision about participants with mixed diagnoses, and "an exercise intervention" needs a decision about interventions where exercise is one component among several.
It helps to write the criteria as a screening form with explicit questions, such as whether the study includes the population, tests the intervention, compares with an eligible comparator, reports an eligible outcome, and has an eligible design. Decision rules for borderline cases should be added as they arise, and recorded with dates so that earlier decisions can be checked for consistency. Changes to criteria made after screening has begun must be reported as deviations from the protocol.
The two screening stages
Title and abstract screening. Reviewers read the title and abstract of each record and decide whether it might meet the criteria. The principle is to be inclusive: if the abstract is unclear or silent on a criterion, move the record to the next stage. This stage removes records that are clearly irrelevant, usually the large majority.
Full-text screening. The full report is retrieved and assessed against every criterion. Each excluded study needs a recorded reason, using the first criterion that fails in a fixed order, so that reasons are not counted twice. Full-text screening also gives the opportunity to link multiple reports of one study. Records for which the full text cannot be found should be listed, and efforts to obtain them, including contacting authors, recorded.
The numbers at each stage feed the PRISMA 2020 flow diagram. A simulated example illustrates the arithmetic: 1,842 records identified from all sources, 612 duplicates removed, leaving 1,230 to screen. After title and abstract screening, 1,150 were excluded and 80 full texts were assessed. Of these, 52 were excluded for the reasons shown, and 28 studies were included.
| Reason | Reports |
|---|---|
| Wrong population | 14 |
| Wrong comparator | 11 |
| Wrong outcome | 9 |
| Wrong design | 10 |
| Conference abstract only, no data | 5 |
| Duplicate report of an included study | 3 |
The reasons sum to the 52 excluded reports, as they must. A flow diagram that does not add up is a common reason for a request to revise.
Dual independent screening
Screening involves judgment, and one person working alone will make errors. Studies of screening show that a single reviewer misses a proportion of eligible records, and that two independent reviewers find more. The Cochrane Handbook expects at least two people at full-text stage. For title and abstract screening, dual screening of all records is the conservative standard, and some reviews use a reduced approach, such as one reviewer with a second checking a sample or the excluded records, particularly in rapid reviews, with the limitation stated.
Reviewers screen independently and without seeing each other's decisions. Conflicts are then resolved by discussion, and a third reviewer arbitrates when they cannot agree. The conflict list is also a source of information about unclear criteria, and may lead to refinement of the screening form.
Measuring agreement. Agreement is often reported with Cohen's kappa, which corrects the observed agreement for the agreement expected by chance. In a simulated pilot of 200 records, both reviewers include 30, reviewer A alone includes 10, reviewer B alone includes 8 and both exclude 152. Observed agreement is 0.910. Chance agreement, from each reviewer's marginal rates, is 0.686. Kappa is (0.910 - 0.686) / (1 - 0.686) = 0.71. Landis and Koch labeled values of 0.61 to 0.80 as substantial agreement, a label that is conventional and not a standard. Kappa behaves oddly when one category is rare, as inclusion usually is, so raw agreement and the number of conflicts are useful companions. The more important check is practical: conflicts should be few enough to resolve, and resolution should produce consistent rules.
Piloting and calibration
Before screening in earnest, all reviewers should screen the same small sample, often 50 to 100 records, then compare results and discuss every disagreement. The aim is to calibrate shared understanding of the criteria, and to revise the form where it caused confusion. A second pilot may be needed if agreement is low. This investment prevents much more work later, and the pilot results can be reported in the methods. Training also applies to anyone new to the team joining part-way through.
Tools and automation
Specialized screening tools allow dual independent screening, blinding, conflict resolution and automatic counts for the flow diagram. Many offer machine-learning prioritization, which ranks records by their likelihood of being relevant, based on the decisions made so far, so that eligible studies are seen earlier. These tools can reduce workload. They do not replace human judgment, and stopping rules based on prioritization are still debated. If a tool with automation is used, report which one, how it was used, what stopping rule applied, and what checks were made. Large language models are being tested for screening, with promising but variable results, and their use in a review requires validation on a sample against human decisions and complete disclosure. The AI-use policy of the journal and of the review team should be followed.
Common mistakes
- Deduplicating once and trusting the result without checking borderline matches.
- Counting multiple reports of one study as separate studies.
- Starting screening with criteria that are not written down.
- Excluding at title and abstract stage on a criterion the abstract does not address.
- Screening by one reviewer without stating this as a limitation.
- Recording no reason for full-text exclusions, or several reasons per record.
- Reporting a flow diagram whose numbers do not add up.
- Reporting kappa without saying at which stage and on what sample.
Practical problems at full text
Other languages. Excluding studies because they are not in a language the team reads introduces bias. Options include asking colleagues, using professional translation for the eligibility decision and extraction, or using machine translation with a check by a speaker of the language. State what was done. If non-English studies were excluded, say so as a limitation.
Full texts that cannot be found. Interlibrary loan, contacting corresponding authors and checking repositories recover many. Those that remain unavailable should be listed as awaiting classification, with the number reported, since they could change the conclusions if they turned out to be large studies.
Conference abstracts and preprints. Decide in the protocol whether they are eligible. If they are, plan how to deal with the limited information they contain, and how to check later whether a full report appeared.
Retracted and corrected papers. Check included studies for retraction notices and expressions of concern, both at screening and again before submission, since status can change while the review is under way.
Keeping an audit trail
A review is open to challenge, from peer reviewers, guideline panels and later authors, and a clear record of decisions is the best defense. Keep the exported search results from each database, the deduplication log, the screening forms and their versions, each reviewer's decisions, the resolution of conflicts, and the list of excluded full texts with reasons. Many review tools keep this automatically. If screening is done in a spreadsheet, lock each version and date it. This record also makes updates far cheaper, because the next team can see what was decided and why, and can screen only the new records against the same criteria. Share the screening form and the list of excluded studies as supplementary material when the journal allows, because open materials let readers check the process. Agree who is responsible for each file and where it lives before the work starts, so that the record is complete when a question arrives months later and the person who made a decision is no longer available to explain it.
How we can help
We can deduplicate records with a documented process, set up screening forms and pilots, carry out dual independent screening, resolve conflicts, link multiple reports and produce the PRISMA flow diagram. [OWNER VERIFICATION REQUIRED] The relevant services are systematic review and literature search strategy.
Frequently asked questions
Do I need two reviewers for screening?
Two independent reviewers is the standard, at least at full text. Reduced approaches exist for rapid reviews, with the limitation stated.
What is Cohen's kappa?
A measure of agreement between two raters that corrects for chance agreement. It is reported alongside raw agreement, and it can be low when one category is rare.
Should I include records with unclear abstracts?
Yes, at the title and abstract stage. Move them to full-text review and decide there.
How should I handle several reports of one study?
Link them, treat the study as the unit of analysis, and extract data from the most complete report while checking the others.
Why record reasons for exclusion?
PRISMA 2020 asks for reasons for excluding full-text reports, which lets readers judge whether the exclusions were appropriate.
Can I use AI tools for screening?
Machine-learning prioritization is common, and large language models are being tested. Validate against human decisions, disclose the use and follow the journal's policy.
References
- Lefebvre C, Glanville J, Briscoe S, et al. Chapter 4: Searching for and selecting studies. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
- Edwards P, Clarke M, DiGuiseppi C, Pratap S, Roberts I, Wentz R. Identification of randomized controlled trials in systematic reviews: accuracy and reliability of screening records. Stat Med. 2002;21(11):1635-1640.
- Waffenschmidt S, Knelangen M, Sieben W, Buhn S, Pieper D. Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review. BMC Med Res Methodol. 2019;19:132.
- Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33(1):159-174.
- Bramer WM, Giustini D, de Jonge GB, Holland L, Bekhuis T. De-duplication of database search results for systematic reviews in EndNote. J Med Libr Assoc. 2016;104(3):240-243.
- Hamel C, Hersi M, Kelly SE, et al. Guidance for using artificial intelligence for title and abstract screening while conducting knowledge syntheses. BMC Med Res Methodol. 2021;21:285.