Guide

Best-evidence synthesis

Best-evidence synthesis keeps the systematic search and explicit criteria of a review but restricts the synthesis to the studies that are most relevant and most credible for the question. It was proposed in education research, where studies vary widely in quality and context. This guide explains the method, compares it with conventional meta-analysis and sets out when it is a reasonable choice.

What best-evidence synthesis is

Best-evidence synthesis is an approach to reviewing research that grew out of dissatisfaction with two extremes. At one end are narrative reviews, which are selective and opaque. At the other end are meta-analyses that combine every study found, regardless of quality or relevance, on the grounds that analysis can adjust for the differences. Robert Slavin argued in the mid-1980s that a review should instead apply to its included studies the same standards of methodological adequacy and relevance that a careful reader would, and should synthesize the best available evidence in the manner of a meta-analysis but with a narrower and better-defined set.

The approach has been used mainly in education, notably in the Best Evidence Encyclopedia and in reviews of reading and mathematics programs, and it has influenced evidence reviews in social policy. It is not as formally standardized as the Cochrane or Campbell methods, and the details vary between authors, which is a source of both flexibility and criticism.

How it differs from a conventional meta-analysis

The two approaches share many steps, including a comprehensive search, clear criteria and the calculation of effect sizes. They differ in the role of study quality and in how the synthesis is carried out.

Conventional systematic review with meta-analysis compared with best-evidence synthesis
FeatureConventional approachBest-evidence synthesis
Selection of studiesAll studies meeting broad criteria, assessed for risk of biasStudies meeting stricter, explicit quality and relevance criteria
Role of study qualityAppraised and explored in analysisUsed to decide inclusion
SynthesisUsually pooled statistically where studies are comparableOften narrative and tabular, with effect sizes, sometimes pooled
AimEstimate the average effect and its variationGive the best available answer for a specific context
Main riskAveraging over flawed or dissimilar studiesSubjective or unclear selection rules

The difference should not be exaggerated. A modern systematic review also considers risk of bias, and often restricts the primary analysis to studies of adequate design. The distinction is a matter of emphasis: best-evidence synthesis puts the quality filter at the door, while conventional reviews often let studies in and examine quality later through sensitivity analysis and subgroups.

Selecting the best evidence

The heart of the method is a set of explicit inclusion standards that go beyond topic. They should be written in advance and justified. In education research, typical criteria concern the following.

  • Design. Randomized or well-matched quasi-experimental comparison, with a control or comparison group.
  • Equivalence at baseline. Evidence that groups were similar before the intervention, or statistical control for pretest differences.
  • Sample size and unit of analysis. A minimum number of classes or schools, not only students, when treatment was assigned to groups.
  • Duration. A minimum length of intervention, to exclude brief laboratory-style studies when the question concerns real use.
  • Outcome measures. Measures that are not inherently tied to the treatment, such as tests made by the developers of the program, and that are of reasonable reliability.
  • Relevance. Settings, populations and interventions similar to those the review is meant to inform.
  • Attrition. Limits on differential or high dropout.

The criteria should come from the question and the field, not from a wish to reach a conclusion. They should be recorded in a protocol, applied by two people independently, and reported with a list of the studies excluded for failing them. In health, the corresponding step is an assessment of risk of bias with a recognized tool, and many authors now apply such a tool to define the best evidence, for example by restricting a primary analysis to trials at low risk of bias.

Synthesizing the studies

Once the set is chosen, each study is described and its effect size calculated, usually as a standardized mean difference, with its confidence interval. The synthesis then combines the information in a way that serves the question.

  • Tables of study characteristics and effect sizes. Each study is described in a standard format, so that the reader can see the evidence directly.
  • Narrative synthesis. A structured account of patterns across studies, taking account of context, implementation and quality. Guidance such as the SWiM reporting guideline for synthesis without meta-analysis applies.
  • Pooled estimates where appropriate. If the included studies are similar enough, a statistical summary, often a weighted mean effect size, adds precision. In education reviews the summary is often reported as a median or weighted mean effect size.
  • Consideration of moderators. Study features, such as grade level, duration or program type, are examined to see whether effects differ.
  • Comparison with the wider set. It is good practice to show how conclusions would change if the excluded studies were added, so that readers can judge the effect of the criteria.

Strengths and criticisms

The method has real strengths. It acknowledges that studies are not equal, and avoids a result that is an average of excellent and poor work. It produces conclusions that are tied to credible evidence and are therefore useful to decision makers, and in fields where many studies are small, brief or confounded, it keeps them from dominating. It also aligns with how practitioners think, which is to ask what the best studies show.

The criticisms are also real. Exclusion criteria involve judgment, and different reviewers can reach different sets and different conclusions. Criteria can be tuned, deliberately or not, to give a particular answer, which is why advance specification matters. Excluding studies discards information, and when the best studies are few, the pooled result can be imprecise. If the criteria favor certain designs, such as short randomized trials, the results may not generalize to routine practice. And a lack of standardization means that two reviews both described as best-evidence synthesis may differ greatly in rigor.

Modern conventional methods address some of the same concerns through risk-of-bias tools, sensitivity analyses restricted to high-quality studies, and GRADE ratings. The differences between the methods have narrowed, and a review can combine elements of both, for example by prespecifying a primary analysis of studies at low risk of bias with sensitivity analyses including all studies.

When it is a reasonable choice

Best-evidence synthesis is most suitable when the literature is large and uneven, when quality varies widely and many studies are weak, and when the users want an answer about effectiveness in realistic settings. It suits education, social programs and policy evaluation, where randomized trials are fewer and quasi-experimental studies vary in rigor. It is less suitable when there are only a few studies, since restricting the set may leave nothing, or when the aim is to describe the whole literature, for which a scoping or evidence-map review is the better choice. In health interventions, a conventional systematic review with risk-of-bias assessment, GRADE and a prespecified sensitivity analysis is the expected standard, and journals may ask authors to use it.

Reporting a best-evidence synthesis

Describe the protocol and registration, the search, the explicit quality and relevance criteria and their justification, the process of applying them, the number of studies excluded at each step and the reasons, and a table of included studies. State how effect sizes were calculated, how they were combined, how moderators were examined, and how the conclusions would change with a broader set of studies. Use a reporting guideline appropriate to the design, such as PRISMA 2020 for the review and SWiM for synthesis without meta-analysis. Be explicit about the limits: the evidence is the best available, which is not the same as strong.

Common mistakes

  • Choosing the inclusion criteria after seeing the results.
  • Failing to report the studies that were excluded and why.
  • Describing a conventional review as best-evidence synthesis without a quality filter.
  • Using criteria so strict that only a handful of studies remain, with no statement of the resulting imprecision.
  • Failing to show what happens when the excluded studies are added.
  • Pooling effect sizes from different outcome types without thought.
  • Ignoring the unit of analysis in cluster-assigned studies.

A worked illustration

Eight simulated studies of an educational program report standardized effect sizes. Four meet the review's prespecified criteria for the best evidence: random assignment or matched comparison, equivalent groups at baseline, an independent outcome measure and adequate duration. Four do not.

Eight simulated studies with effect size, standard error and whether they meet the best-evidence criteria
StudyEffect sizeStandard errorMeets criteria
Study A0.180.06Yes
Study B0.220.09Yes
Study C0.120.08Yes
Study D0.280.12Yes
Study E0.550.20No
Study F0.700.25No
Study G0.480.18No
Study H0.620.22No

The weighted mean effect size of the four best studies is 0.18 (standard error 0.04). With all eight studies, the weighted mean is 0.23 (standard error 0.04). The studies that fail the criteria are the ones with the larger effects, which is a common pattern: weaker designs, such as unmatched comparisons and developer-made tests, tend to overstate benefit. A review that pooled everything would have reported an effect that is about 27 percent larger than that from the best studies. In this example the narrower set is also less precise, as should be expected when studies are discarded. Reporting both figures lets the reader see the effect of the criteria. The data are simulated. Note also what the example does not show: it does not prove that the weaker studies are wrong, since they may have been run in different conditions, only that their results differ in a direction that known biases would predict. That is why the narrative should discuss both, and why the choice of criteria needs a defense based on evidence about bias in the field, such as studies comparing developer-made with independent tests, and not simply on preference.

How we can help

We can set up explicit quality and relevance criteria, search and screen studies, calculate and synthesize effect sizes, run sensitivity analyses with the broader set and report the work to PRISMA 2020 and SWiM. [OWNER VERIFICATION REQUIRED] The relevant services are systematic review and meta-analysis.

Frequently asked questions

What is best-evidence synthesis?

A review approach that applies explicit methodological and relevance criteria to select the strongest studies, then synthesizes them with effect sizes and narrative, sometimes with pooling.

Who proposed it?

Robert Slavin, in the mid-1980s, mainly for educational research.

How is it different from meta-analysis?

Conventional meta-analysis usually includes all eligible studies and examines quality later. Best-evidence synthesis uses quality and relevance criteria to select studies first.

Is it less rigorous?

Not if the criteria are explicit, prespecified and reported. The risk is subjective or post hoc criteria.

When should I use it?

When the literature is large and varies widely in quality, and users want conclusions based on credible studies. It is common in education and social policy.

Should I show results with the excluded studies?

Yes. A sensitivity analysis with a broader set shows how much the criteria affect the conclusions.

References

  1. Slavin RE. Best-evidence synthesis: an alternative to meta-analytic and traditional reviews. Educ Res. 1986;15(9):5-11.
  2. Slavin RE. Best evidence synthesis: an intelligent alternative to meta-analysis. J Clin Epidemiol. 1995;48(1):9-18.
  3. Campbell M, McKenzie JE, Sowden A, et al. Synthesis without meta-analysis (SWiM) in systematic reviews: reporting guideline. BMJ. 2020;368:l6890.
  4. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71
  5. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  6. Grant MJ, Booth A. A typology of reviews: an analysis of 14 review types and associated methodologies. Health Info Libr J. 2009;26(2):91-108.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.