Evidence synthesis in public policy
Public policy research asks what policies do: whether a tax change alters behavior, whether a job program raises employment, whether a regulation improves safety, whether an education reform changes outcomes, and who gains or loses. Evidence synthesis supports decision-makers who cannot read every study, and networks such as the Campbell Collaboration and what-works centers produce reviews for policy audiences. Syntheses come in several forms: full systematic reviews, rapid reviews, evidence and gap maps, meta-analyses and realist syntheses.
The field presents particular challenges. Many policies cannot be randomized. Effects depend on context, implementation and population. Decisions are time-pressured. The people who commission a review may have views about the answer. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for social sciences. The service provides research synthesis only and does not make policy recommendations.
Choosing the form of synthesis
| Form | Best suited to | Main limits |
|---|---|---|
| Systematic review with meta-analysis | Well-defined interventions with comparable outcomes | Needs enough similar studies; heterogeneity can be high |
| Evidence and gap map | Showing where evidence exists and where it is missing | Describes studies; does not estimate effects |
| Rapid review | Time-limited decisions | Shortcuts increase risk of missing studies or errors; shortcuts must be stated |
| Realist synthesis | Understanding how and why programs work in context | Interpretive; not for estimating average effects |
| Mixed-methods synthesis | Combining effects with implementation and experience | Demanding; integration methods vary |
The choice depends on the question and the time available. A common mistake is to commission a pooled estimate when the evidence is thin or too varied, and a map or a realist synthesis would serve the decision better. A review should state why it used its chosen form. Our pages on rapid review and realist review describe two of these in detail.
Quasi-experimental designs and credibility
Policy evaluation often relies on natural experiments. Difference-in-differences compares changes over time between affected and unaffected groups and assumes parallel trends. Regression discontinuity uses eligibility thresholds. Synthetic control builds a comparison from a weighted combination of unaffected units. Interrupted time series examines changes in level and trend after a policy. Each has assumptions that can be checked in part, and a review should assess whether the primary studies tested them, for example with pre-trend plots or placebo tests. Estimates from these designs apply to specific groups: the local effect near a threshold, or the effect on the units that were treated, which may differ from the effect elsewhere.
When studies use different designs for the same policy, results can be reported by design. Studies using aggregate data with few units often have large uncertainty that standard errors understate if they ignore correlation over time or within regions.
Implementation, context and transferability
A policy that works in one place may fail in another because of different institutions, resources, capacity or population. Evidence on implementation, including fidelity, dosage and local adaptation, is as important as effect estimates. A review should extract descriptions of implementation, and where possible link them to outcomes. Frameworks that describe context, such as those used in realist and implementation research, help to organize information. Transferability is judged by comparing the conditions in the studies with those in the target setting, which requires information that a review of studies cannot supply on its own. We provide the structured information needed for such judgments and do not make them.
Distributional effects and equity
Average effects can hide differences across groups, and a policy that raises the average can widen gaps. Reviews should extract effects by income, sex, race and ethnicity, age and region where studies report them, and test group differences as exploratory analyses. Subgroup analyses in single studies are often underpowered, and absence of reported subgroup data is itself a finding. Evidence on equity is thin for many policies, and a synthesis should say so. Use of language about groups should be careful, precise and consistent with the sources.
Communicating certainty and uncertainty
Decision-makers need to know how sure we are. Frameworks such as GRADE and others designed for public health and social policy rate the certainty of evidence, taking account of design, consistency, directness, precision and bias. A review should present effect estimates with intervals and prediction intervals, report the certainty rating with reasons, and state plainly what remains unknown. Overstating confidence erodes trust in evidence over time, and so does a bland statement that more research is needed. Plain-language summaries should keep numbers and avoid terms like "proven" or "disproven", which overstate what a set of studies can show.
Costs and economic evaluation
Policy decisions need information on cost as well as effect. Many evaluations do not report costs, or report them inconsistently. A review can extract cost data where available and standardize currencies and years, as described on the health economics page, but should note when cost information is missing, which is common. Cost-effectiveness depends strongly on local prices and scale, so transferring results needs local modelling.
Independence and stakeholder involvement
Reviews commissioned by agencies may be shaped by the agency's interests. Good practice includes a protocol agreed before searching, transparent inclusion rules, disclosure of funding and conflicts of interest, and engagement with stakeholders on the question and the interpretation, without allowing them to select the studies or edit the findings. We follow these practices and record the role of each party in the report.
Searching for policy evidence
Policy evidence is spread across academic journals, working-paper series, government evaluations, think-tank reports and international organizations. Much of the best evaluation work is published first as a report, and sometimes never in a journal. Searches that stop at bibliographic databases will miss it. We combine database searches (Scopus, Web of Science, EconLit, Sociological Abstracts, ERIC, PAIS) with searches of repositories, targeted websites of agencies, citation tracking from known studies and contact with experts. We document each source, the date and the search terms, and we record how many records came from each source so that readers can see how much depends on grey literature. Language restrictions and country coverage are stated, since policy evidence from non-English-speaking countries is often under-represented.
Search results are screened by two reviewers on a sample, with the agreement reported and the screening criteria refined if disagreement is high. Reports that describe the same evaluation in several documents are linked to a single study record to avoid double counting.
Timeliness and living evidence
Policies change faster than evidence reviews can be completed. Living reviews, which are updated as new studies appear, are one response, and they are useful when a decision is pending and new studies arrive often. They require a plan for updating, a rule for when to change conclusions and resources to maintain the work. Another response is a rapid review with stated shortcuts. In either case, the review should show the date of the last search and state whether the findings could be affected by recent events, such as a change in the law or in economic conditions, that studies from earlier periods do not capture, and it should describe how the conclusions would change if those events altered the conditions under which the studied policies worked.
Common pitfalls we look for
- Pooling studies with very different policies under one label.
- Treating a local or threshold effect as universal.
- Ignoring implementation and context.
- Reporting only average effects when distribution matters.
- Not stating how rapid-review shortcuts affect confidence.
- Selecting studies to match a preferred conclusion.
Planning a policy synthesis
We help refine the question with the commissioning team, choose the form of synthesis, plan searches across economics, political science, public administration, health and education databases and grey-literature sources (government, think tanks, evaluation repositories), and set up coding of policy features, design, context, implementation, outcomes, equity data and cost. See the systematic review service for scope and process.
An invented example of a policy effect
Suppose a review of a wage subsidy finds an average increase in the probability of employment after one year of 3 percentage points, based on 15 invented quasi-experimental studies, with a 95 percent confidence interval from 1 to 5 points and a prediction interval from minus 2 to 8 points. In a target group where employment probability without the subsidy is 40 percent, 3 points is a relative increase of 7.5 percent. The prediction interval says that in some places the subsidy could do nothing. If studies in recessions show 5 points and those in strong labor markets 1 point, the economic context moderates the effect. The review would report these numbers and the certainty rating, and leave the decision to those responsible.
Coding and transparency
Coding frames record the policy and its components, jurisdiction, level of government, period, target group, design and identification strategy, outcomes and measurement, follow-up, implementation features, costs, funder and evaluator. Two coders work independently on a sample, agreement is reported, and the coded data and code are shared. Registration of the protocol is encouraged, and we note when a rapid approach is used and which steps were limited.
How we support research projects in this area
From a policy or practice question to a published review
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
Help choosing the review type, a protocol, and a plan for equity and context, with registration where suitable.
Searching and extraction
Searches including grey literature, screening, and extraction for quantitative and qualitative strands.
Synthesis
Meta-analysis, qualitative or mixed-methods synthesis, or an evidence map, as the question requires.
Manuscript and submission
The report or manuscript, a policy summary, and journal or funder preparation.
Boundaries of this service
A public policy synthesis describes published evidence on the effects of policies and programs. It does not make policy recommendations, give political or lobbying advice, or provide legal opinions. Effects depend on context and implementation, and evidence may not transfer to a different jurisdiction or population.
Frequently asked questions
What is the difference between a systematic review and an evidence map?
A review estimates or synthesizes effects, whereas a map shows what research exists and where gaps are, without estimating effects.
Why not always do a rapid review?
Shortcuts raise the chance of missing studies or errors. They are acceptable for time-limited decisions if stated clearly.
Can quasi-experimental studies be combined?
Where outcomes and designs are comparable, with separate analyses by design and attention to the assumptions.
How do you address context?
By extracting implementation and setting details, using moderators with caution and describing what is known and not known about transfer.
How is certainty communicated?
With effect estimates, intervals, prediction intervals and a certainty rating with reasons.
Do you make policy recommendations?
No. The service provides research and evidence-synthesis support only.
References
- Petticrew M, Roberts H. Systematic reviews in the social sciences: a practical guide. Oxford: Blackwell; 2006.
- Snilstveit B, Vojtkova M, Bhavsar A, Stevenson J, Gaarder M. Evidence & gap maps: a tool for promoting evidence informed policy and strategic research agendas. J Clin Epidemiol. 2016;79:120-129.
- Pawson R, Greenhalgh T, Harvey G, Walshe K. Realist review: a new method of systematic review designed for complex policy interventions. J Health Serv Res Policy. 2005;10(Suppl 1):21-34.
- Craig P, Dieppe P, Macintyre S, Michie S, Nazareth I, Petticrew M. Developing and evaluating complex interventions: the new Medical Research Council guidance. BMJ. 2008;337:a1655.
- Hamel C, Michaud A, Thuku M, et al. Defining rapid reviews: a systematic scoping review and thematic analysis of definitions and defining characteristics of rapid reviews. J Clin Epidemiol. 2021;129:74-85.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.