Meta-analysis and evidence synthesis for development economics

Development economics studies poverty, growth and the effects of programs and policies in low- and middle-income settings. Its evidence includes randomized impact evaluations, quasi-experiments and cross-country regressions. Reviews have to deal with outcomes on many scales, programs that differ in design and setting, small numbers of trials per intervention and the question of whether results travel.

Evidence synthesis in development economics

Development economics asks which programs and policies reduce poverty and improve well-being in low- and middle-income countries. Over recent decades the field has built a large body of randomized evaluations of cash transfers, microcredit, school interventions, health programs and agricultural support, alongside quasi-experimental studies and cross-country growth regressions. Systematic reviews and meta-analyses are used by governments, donors and researchers to summarize what the evidence shows about a given type of program.

Synthesis in this field raises issues that differ from clinical reviews. Programs labelled alike are often delivered differently. Outcomes are measured on many scales and over varying time horizons. The number of studies of a given intervention is often small, and each is run in a specific place at a specific time. Whether the average effect describes what would happen elsewhere is the main question rather than a side issue. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for economics.

Designs and their credibility

Common designs in impact evaluation
DesignIdentifying assumptionIssue for synthesis
Cluster randomized trialRandom assignment of villages, schools or clinicsClustering must be accounted for; spillovers between units; attrition
Difference-in-differencesParallel trends absent the programAssumption cannot be fully tested; timing of treatment may vary
Regression discontinuityUnits just either side of a cutoff are comparableLocal effect near the cutoff only; bandwidth choices
Instrumental variablesInstrument affects outcome only through treatmentEstimates a local effect for compliers; instrument strength
Cross-country regressionControls capture confoundingFew observations, many possible models; weak causal inference

A review should treat designs separately or use design as a moderator, and appraise each with a suitable tool, such as a risk-of-bias assessment for randomized trials adapted to clusters, and tools for non-randomized studies. Attrition, non-compliance and spillovers are common issues. Cross-country regressions are particularly hard to synthesize, since the results depend on model specification and a small number of countries, and reviews of this literature should treat them as a distinct line of evidence.

Outcomes on many scales

Outcomes include consumption in local currency, test scores on different tests, enrolment rates, anthropometric indicators and indices of empowerment. A common approach is to convert continuous outcomes to standardized mean differences, and monetary outcomes to a percentage change relative to the control mean. Standardized effects depend on the spread in the sample, which varies across settings, so a unit of standard deviation does not mean the same everywhere. Percentage changes depend on the baseline. A review should state its choice, give some effects in natural units, and not rely on a single standardization.

Several evaluations report many outcomes, often in indexes. Studies that test many outcomes without adjusting for multiple comparisons raise concern about false positives. We extract pre-specified primary outcomes when a registry or plan exists, and note when outcomes were selected after the data were seen.

Heterogeneity and external validity

The central question for development evidence is how far an effect found in one setting predicts the effect in another. Bayesian hierarchical models estimate the average effect and the variation across sites, and they allow the effect in a new site to be described with its predictive distribution. Studies of the cash transfer and microcredit literature have found average effects with considerable variation across sites, and a review should present the between-study variation clearly, not only the mean. Meta-regression can examine moderators such as baseline poverty, program design features and implementer type, but the number of independent sites limits this.

A well-documented concern is that programs tested in trials are often run by motivated organizations with extra resources, so effects at scale may be smaller. A review can compare small pilot studies with larger government-run programs where both exist, and should say when this comparison is not possible.

Cash transfers, microfinance and livelihoods

Reviews of cash transfers have examined effects on consumption, school attendance, health service use and employment, with attention to whether transfers are conditional on behavior. Reviews of microcredit have found modest average effects on business activity and little effect on household income or poverty in several trials. Graduation programs, which combine asset transfers with training and support, have shown more promising effects in some trials. These conclusions are about the programs and settings studied, and a review should resist stating that an intervention works or does not work in general. The honest summary describes the average effect, its range and the contexts covered.

Education and health programs

Education studies measure enrolment, attendance and learning. Learning outcomes are measured by tests that differ by country, and standardizing within each study is common. Health program evaluations measure service use, nutrition, infection or mortality, and some of these outcomes need long follow-up. Reviews of such programs should consider the time horizon, the cost where reported, and the potential for spillovers across treated and untreated units. We link to relevant clinical reviews in the medical and health sciences area when health outcomes are the focus.

Dependent effects and small numbers of studies

Evaluations often report several outcomes and several comparison arms, and multiple papers may report on the same trial at different follow-up times. Robust variance estimation and multilevel models handle this, but both need enough studies to be reliable. When studies are few, as is common, we favor simpler models, avoid overfitting moderators and report the uncertainty of the heterogeneity estimate. Bayesian models with sensible priors can be helpful when data are sparse, with the prior stated and sensitivity analyses shown.

Publication bias and registration

Trial registration is much less complete in development economics than in medicine, but registries exist, such as the AEA RCT Registry, and a review can compare registered with published outcomes. Studies with significant results are more likely to be published in journals, and working papers may report different results from later versions. We search working-paper series, institutional repositories and evaluation databases, and compare versions when more than one exists. Funnel-plot and selection-model methods are applied with their limits explained.

Spillovers, attrition and follow-up

Three features of development trials affect how results should be read. First, spillovers: a program given to some households in a village can change prices, behavior or knowledge for neighbors who did not receive it, so the comparison between treated and control units may understate or overstate the effect. Studies that randomize at a higher level, such as the village, handle this better, and a review should note the level of randomization. Second, attrition: if people leave the sample in ways related to the program, estimates are biased, and studies report attrition rates and bounds of different quality. Third, follow-up length: a short follow-up can miss slow benefits or fading effects, and a review should show effects by time since the program ended wherever data allow, together with the number of studies behind each time point.

Common pitfalls we look for

  • Treating programs with the same label as identical.
  • Pooling randomized and non-randomized studies without separating designs.
  • Presenting an average effect without its range across sites.
  • Ignoring clustering and attrition in primary studies.
  • Selecting outcomes after seeing results.
  • Generalizing from pilots to national programs.

Planning a development economics synthesis

We help define the intervention, comparison and outcomes, decide how designs are handled, plan searches across economics, public health, education and development databases, and set up coding of program design, setting, scale and implementer. Evidence-gap maps can show where studies exist and where they do not, which is helpful when the number of studies per intervention is small. See the systematic review service for scope and process.

Reading an average effect across sites

Suppose a review finds an average effect of 0.10 standard deviations on test scores across eight invented trials, with a between-study standard deviation of 0.08. Roughly, the true effect in a new site might be expected to fall between about minus 0.06 and plus 0.26 (the mean plus or minus two between-study standard deviations, ignoring the uncertainty in the mean itself). The average suggests a small benefit, but a decision-maker should see that a similar program in a new place could do nothing, or much more. Characteristics of the sites and the program, if coded, may explain part of the spread, and the review should describe what is and is not explained.

This is why reporting the range matters as much as the average, and why we do not summarize such a result as "the program works".

Data, ethics and local knowledge

Impact evaluations involve people in low-income settings, and the studies used in a review should have had ethical approval and consent procedures. The reviewer extracts only aggregated results. Local context affects interpretation, and where possible reviews benefit from involving researchers or practitioners from the countries studied when interpreting heterogeneity. Language limits are a concern since evaluations may be published in languages other than English, and the search strategy should state how these were handled.

How we support research projects in this area

Support

From estimates to a published meta-regression

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A question, the effect size and conversion rules, the plan for specification variables and a protocol in line with MAER-Net guidelines.

  • Searching and extraction

    Searches of EconLit, Scopus, RePEc and working paper series, with coding of estimates and specifications by two people.

  • Synthesis

    Meta-regression with clustered or multilevel models, FAT-PET-PEESE and model-averaging sensitivity analyses.

  • Manuscript and submission

    The manuscript, replication data and code, and journal preparation.

Get a quoteDescribe the parameter, the literature and your target journal.

Boundaries of this service

A development economics synthesis describes average effects and their variation across the studies it includes. It does not design programs, advise governments or donors on allocation, or predict what a program will do in a particular community. Evidence comes from specific settings and may not transfer, and outcomes measured in studies may not capture everything that matters to the people affected.

Frequently asked questions

Can randomized trials of development programs be meta-analyzed?

Yes, when interventions and outcomes are comparable enough. The review presents the average effect together with how much it varies across sites.

How do you deal with outcomes on different scales?

By standardizing or using percentage changes, stating the choice and showing results in natural units where possible.

Why not just report the average effect?

Because effects differ across settings, and the range is what tells a reader how well the average might apply elsewhere.

Are non-randomized studies included?

Where relevant, but separately or with design as a moderator, and appraised with suitable tools.

Do results scale up from pilots?

Not always. We compare small pilots with larger programs when both exist, and say when we cannot.

Do you advise on aid or program design?

No. The service provides research and evidence-synthesis support only.

References

  1. Vivalt E. How much can we generalize from impact evaluations? J Eur Econ Assoc. 2020;18(6):3045-3089.
  2. Meager R. Understanding the average impact of microcredit expansions: a Bayesian hierarchical analysis of seven randomized experiments. Am Econ J Appl Econ. 2019;11(1):57-91.
  3. Banerjee A, Karlan D, Zinman J. Six randomized evaluations of microcredit: introduction and further steps. Am Econ J Appl Econ. 2015;7(1):1-21.
  4. Banerjee A, Duflo E, Goldberg N, et al. A multifaceted program causes lasting progress for the very poor: evidence from six countries. Science. 2015;348(6236):1260799.
  5. Waddington H, White H, Snilstveit B, et al. How to do a good systematic review of effects in international development: a tool kit. J Dev Eff. 2012;4(3):359-387.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.