Why one database is not enough
Databases index different journals, in different languages, with different indexing practices and different lags. Studies of overlap show that no single database captures every eligible trial or study, and that searching several raises the number found. Relying on one source is a common weakness of reviews that are labelled systematic, and peer reviewers and tools that appraise reviews, such as AMSTAR 2, look for searching in at least two sources along with other supplementary methods.
The aim of the search is not to retrieve everything that exists, which is not achievable, but to find a set of studies that is complete enough that missing studies are unlikely to change the conclusions. Reaching that depends on the choice of sources as much as on the quality of the search terms. A good strategy run in the wrong database finds little, and a mediocre one run in the right databases finds more.
Core databases by discipline
The table lists databases that are commonly used. It is not a complete list, and a librarian or information specialist is the best guide for a specific topic. Coverage, interfaces and access arrangements change, so check current details before relying on them.
| Database | Main coverage | Notes |
|---|---|---|
| MEDLINE (via PubMed or Ovid) | Biomedicine and health, including nursing and dentistry | Controlled vocabulary (MeSH); free through PubMed |
| Embase | Biomedicine with strong drug and device coverage and European journals | Controlled vocabulary (Emtree); subscription |
| CENTRAL (Cochrane Central Register of Controlled Trials) | Reports of controlled trials, from several sources | Part of the Cochrane Library |
| CINAHL | Nursing and allied health | Subscription; own subject headings |
| PsycINFO | Psychology and behavioral sciences | Subscription; Thesaurus of Psychological Index Terms |
| ERIC | Education | Free; own thesaurus |
| EconLit | Economics | Subscription; JEL classification codes |
| Scopus and Web of Science | Multidisciplinary citation databases | Subscription; citation searching |
| Global Index Medicus / regional databases | Journals from low- and middle-income regions (for example LILACS, African Index Medicus, IMEMR) | Free; coverage and interface vary |
| Specialty databases | For example SPORTDiscus, ProQuest Dissertations, ASSIA, AGRICOLA, GreenFILE | Depends on topic |
For a review of health interventions, the Cochrane Handbook recommends searching CENTRAL, MEDLINE and Embase as a minimum where possible, and adding regional and subject databases according to the topic. For psychology, education, management and economics, PsycINFO, ERIC, Business Source Complete, EconLit and discipline-specific sources carry the core of the literature, and the multidisciplinary citation databases help with fields where journals are not well covered by the health databases.
How to choose for your topic
Work through these questions before writing the strategy.
- Which fields publish studies on this topic? An intervention for adolescent depression is in the health databases and in education and psychology sources. An environmental exposure may appear in toxicology, public health and ecology.
- Which study designs are eligible? Trials are well captured by CENTRAL and the trial registries. Observational studies, qualitative studies and economic evaluations need different sources and filters.
- Which languages and regions matter? If the condition is concentrated in particular regions, regional databases can add studies that international databases miss.
- What type of review is it? A rapid review may justify fewer sources with a stated limitation. A guideline-supporting systematic review or a Cochrane-style review expects wide coverage.
- What resources are available? Subscription access is a practical constraint. An institutional library or a librarian can often provide access that individuals lack.
Write the reasons for the choice in the protocol, and ask a librarian to review the search, as the PRESS checklist describes. Peer review of search strategies finds errors that change what is retrieved.
Beyond bibliographic databases
Databases do not index everything, and several complementary methods are standard.
- Trial registries. ClinicalTrials.gov, the WHO International Clinical Trials Registry Platform and national registries identify completed trials, some of which have never been published, and show the outcomes that were planned. Registry searching also helps to assess selective outcome reporting.
- Citation chasing. Checking the reference lists of included studies and relevant reviews (backward chasing), and finding the papers that cite them (forward chasing), through Scopus, Web of Science or Google Scholar, finds studies that keyword searches miss.
- Grey literature. Dissertations, conference abstracts, reports, regulatory documents and preprints. The separate guide on grey literature gives detail.
- Hand-searching key journals and conference proceedings, when the field is concentrated.
- Contact with experts and authors to ask about unpublished or ongoing studies.
- Regulatory sources such as drug approval documents, which can contain trial data not found in journals.
Google Scholar is useful for citation chasing and for finding grey literature, but it is not suited as a primary database for systematic searching, because its search syntax is limited, its results cannot be exported in bulk reliably and the exact content is not transparent. If it is used, screen a stated number of the top results and report that.
Translating a strategy between databases
A strategy is developed in one database, usually MEDLINE, and then translated for the others. Translation is not copying. Controlled vocabularies differ, so Medical Subject Headings must be mapped to Emtree terms in Embase or to thesaurus terms in PsycINFO. Field codes, truncation symbols, proximity operators and how phrases are treated differ between platforms. A strategy that works in PubMed can return nonsense in Ovid or fail without an error message. Each translated search should be tested by checking that known relevant studies are retrieved, and the final strategy for each database should be saved exactly as run.
Search filters, which are prebuilt blocks of terms for study designs such as randomized trials, can save time. They should come from validated sources and be applied in the database they were designed for.
Documenting the search
PRISMA 2020 and its search extension, PRISMA-S, ask for the name of each database and its interface, the date of the last search, the full strategy for at least one database and, preferably, every database, the limits applied, and the sources beyond databases. This allows others to repeat and judge the search. A search that cannot be reproduced cannot be assessed.
Record the number of records from each source before deduplication, so that the flow diagram is accurate. Save search histories and export files. Set a date for rerunning the search, usually shortly before submission, because reviews take months and the literature moves on. Many journals and guidelines expect searches no more than a year old at the time of submission, though the expectation varies and should be checked for the target journal.
Date, language and format limits
Limits applied to a search can exclude relevant studies. Restricting by language is a common source of bias, because studies with positive findings are more likely to be published in English, and excluding other languages can affect results in some topics. Restricting by date should be justified by the history of the question, for example the year an intervention became available, and not by convenience. Excluding conference abstracts, preprints or non-peer-reviewed material reduces completeness and should be stated and justified. Where limits are used, describe them in the methods and discuss them in the limitations.
Common mistakes
- Searching a single database, often PubMed alone, and calling the review systematic.
- Choosing sources that fit the author's access rather than the topic.
- Copying a strategy into other databases without adapting the vocabulary and syntax.
- Omitting trial registries, which leaves out unpublished trials and the planned outcomes.
- Failing to record dates, interfaces and full strategies.
- Applying language or date limits without justification.
- Not rerunning the search before submission.
- Relying on Google Scholar as the only supplement.
Examples of source sets for different topics
These sets are illustrations of how the choice works, not prescriptions.
- A drug trial review. MEDLINE, Embase and CENTRAL, plus ClinicalTrials.gov, the WHO registry platform and regulatory documents, with citation chasing from included trials and relevant reviews.
- A school-based intervention. ERIC, PsycINFO, MEDLINE and a multidisciplinary database such as Scopus, plus the What Works Clearinghouse or equivalent sources and dissertation databases.
- A workplace or management question. Business Source Complete, PsycINFO, Scopus or Web of Science and discipline repositories such as SSRN, with attention to working papers, which are the usual first form of publication in economics and management.
- A nursing practice question. CINAHL, MEDLINE, Embase and PsycINFO, with a search for guidelines and quality-improvement reports where relevant.
- An environmental exposure. MEDLINE, Embase, Scopus, Web of Science and Greenfile, with agency reports and theses.
Checking whether the search is complete
There are practical ways to test a search before relying on it. Assemble a small set of studies already known to be eligible, from earlier reviews, expert suggestions or a preliminary scoping search, and confirm that the strategy retrieves each of them in the databases that should index them. If a known study is missed, find out why: a missing synonym, an indexing term that was not used, or a database that does not hold it. After screening, look at where the included studies were found. If most were identified by a single source, or by citation chasing alone, consider whether the other sources were adding anything. If a number of included studies came only from supplementary methods, the main strategy may have been too narrow, and the finding is worth reporting in the limitations. Such checks cost little and make the methods section more credible. Record the outcome too: a line in the methods stating that a set of known studies was used to test the strategy, and how many were retrieved, tells readers that the search was validated and not simply run once.
How we can help
We can select databases suited to your topic and review type, develop and translate search strategies, search registries and grey sources, run citation chasing, document everything to PRISMA-S and rerun searches before submission. [OWNER VERIFICATION REQUIRED] The relevant services are literature search strategy and systematic review.
Frequently asked questions
How many databases should I search?
At least two or three suited to the topic. For health interventions, MEDLINE, Embase and CENTRAL are the usual minimum where access allows, with regional and subject databases added as needed.
Is PubMed enough?
Not on its own. PubMed gives access to MEDLINE and other content, but it does not cover everything that Embase or discipline databases do, and it does not replace registries or citation chasing.
Should I use Google Scholar?
It is useful for citation chasing and grey literature, but it is not suited as a primary source because of limited syntax and export. If used, report how many results were screened.
Do I need to search trial registries?
For reviews of interventions, yes. Registries identify unpublished trials and show planned outcomes, which helps to detect selective reporting.
Which interface should I record?
The platform through which the database was searched, for example Ovid or PubMed, along with the date and the exact strategy.
How current must the search be?
Rerun it shortly before submission. Many journals expect searches within about a year of submission, but check the target journal.
References
- Lefebvre C, Glanville J, Briscoe S, et al. Chapter 4: Searching for and selecting studies. In: Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current edition.
- Rethlefsen ML, Kirtley S, Waffenschmidt S, et al. PRISMA-S: an extension to the PRISMA statement for reporting literature searches in systematic reviews. Syst Rev. 2021;10:39.
- McGowan J, Sampson M, Salzwedel DM, Cogo E, Foerster V, Lefebvre C. PRESS Peer Review of Electronic Search Strategies: 2015 guideline statement. J Clin Epidemiol. 2016;75:40-46.
- Bramer WM, Rethlefsen ML, Kleijnen J, Franco OH. Optimal database combinations for literature searches in systematic reviews: a prospective exploratory study. Syst Rev. 2017;6:245.
- Shea BJ, Reeves BC, Wells G, et al. AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ. 2017;358:j4008.
- Gusenbauer M, Haddaway NR. Which academic search systems are suitable for systematic reviews or meta-analyses? Evaluating retrieval qualities of Google Scholar, PubMed, and 26 other resources. Res Synth Methods. 2020;11(2):181-217.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71