Evidence synthesis in economics
Economics has its own tradition of research synthesis. Stanley and Jarrell introduced meta-regression analysis to the discipline in 1989, and the Meta-Analysis of Economics Research Network (MAER-Net) has since set out standards and held regular colloquia. The subject matter explains the difference from clinical practice. Instead of one treatment effect measured in comparable trials, economics has many estimates of a parameter, each produced by a different model on different data. A study may report dozens of estimates from alternative specifications, and the choice among them is part of the research process.
Three features follow. First, heterogeneity is large and expected, and its sources (data, country, period, method, control variables) are the main object of study. Second, estimates are strongly dependent, since the same data set and the same authors produce many. Third, the incentive to report results that are significant or that agree with theory has been shown to bias estimates in many literatures. Ioannidis, Stanley and Doucouliagos found that in a large sample of economics research most estimates came from studies with low statistical power, with a median power around 18 percent in the sample they examined.
The approach to economic reviews is therefore different from the one for trials, and this page explains it. Related methods are described under meta-regression and the meta-regression service.
Effect sizes in economics
| Measure | Description | Points to check |
|---|---|---|
| Elasticity | Percentage change in one variable for a percentage change in another | Different functional forms give different elasticities; conversions need the means |
| Partial correlation | Association between two variables, holding other regressors constant, computed from t-statistics and degrees of freedom | Allows pooling across specifications; interpretation differs from units-based effects |
| Semi-elasticity | Percentage change in the outcome for a one-unit change in the regressor | Common for returns to schooling; units must match |
| Treatment effect | Estimated effect from an experiment or quasi-experiment | Local average effects differ from average effects |
| Standardized coefficient | Coefficient in standard-deviation units | Depends on the sample variance, which differs across studies |
Economic estimates depend on units, so pooling requires a common effect size. A frequent choice is the partial correlation, because it can be computed from the reported t-statistic and degrees of freedom even when units differ. Its drawback is that it gives a unit-free measure that is harder to read in economic terms, so reviews often convert back to an elasticity or a marginal effect at a reference point. The conversion method should be stated and tested in a sensitivity analysis.
Publication selection and FAT-PET-PEESE
When a literature favors significant or theory-consistent results, estimates are selected. The standard test regresses the estimate on its standard error. In the absence of selection, the estimate should be unrelated to its standard error. A significant slope indicates funnel asymmetry, tested as the funnel-asymmetry test (FAT). The intercept, when the slope is controlled, estimates the effect corrected for selection and is the precision-effect test (PET). The precision-effect estimate with standard error (PEESE) uses the variance instead and is preferred when a genuine effect exists. Stanley and Doucouliagos described this family of methods in 2014. Simulation work suggests the methods perform well under some conditions and poorly under others, notably with heterogeneity, few studies and strongly dependent estimates, so they are best used as sensitivity analyses.
Because of the dependence, standard errors in the regression must be clustered by study, or a multilevel model must be used. Weights matter: unweighted, inverse-variance and inverse-number-of-estimates weights give different answers, and the choice should be justified and tested. Alternatives such as selection models, the p-curve and the weighted average of the adequately powered estimates (WAAP) are also used, and a good review reports several.
Explaining heterogeneity with moderators
The main purpose of a meta-regression in economics is to explain why estimates differ. The dependent variable is the estimate. The regressors describe how it was produced: the country and period, the data type (cross-section, panel, time series), the estimation method (OLS, instrumental variables, fixed effects), the control variables included, the publication outcome and the quality of the journal. With dozens of candidate moderators and limited numbers of studies, the choice of variables is hard, and model uncertainty is large.
Bayesian model averaging is a response: it averages over many model specifications, weighting by their fit, and reports the probability that each variable should be included. This avoids reporting a single, selected specification. Frequentist model averaging and the general-to-specific procedure are alternatives. In any case, the review should state how the set of moderators was chosen, report the results of a model with all of them and of models with fewer, and report the variance inflation or the correlation between moderators.
Estimation quality matters in this literature. Studies using instrumental variables or randomization often report larger standard errors and different estimates than ordinary least squares studies, and a review can test whether identification strategy moderates the result. Authors can then say how much of the variation between studies is due to methods and how much to real differences between contexts.
Reporting guidelines for economics
Stanley and colleagues set out reporting guidelines for meta-analysis in economics in 2013, and Havranek and colleagues updated them in 2020. The guidelines ask authors to say why the meta-analysis was done, to describe the search (databases, working paper series, key words, the date and the number of studies found and rejected), to describe how data were coded and checked by a second person, to describe the effect size and its derivation, to report the sensitivity of results to specification and to make the data and code available. They apply in the same spirit as PRISMA in health, and our reports follow them.
The search in economics uses EconLit, Scopus, Web of Science and RePEc, and working paper repositories such as SSRN and NBER, since many results appear there first. Restricting to journal articles raises the risk of publication selection, and a comparison of published and unpublished estimates is a standard check.
A worked reading of a funnel-asymmetry test
Suppose a review of price elasticities of demand collects 300 estimates from 60 studies. The numbers are invented to show the reasoning and are not from a real review. The unweighted mean is -1.6. A regression of the estimate on its standard error, with standard errors clustered by study, gives a slope that is significant and a PET intercept of -0.7.
The reading is as follows. The slope shows that less precise estimates are more negative, which is consistent with the selection of large negative values, a sign of publication selection. The intercept suggests that after correcting for selection the typical elasticity is nearer to -0.7 than to -1.6. This is a large difference. A careful review would then report the PEESE estimate, the WAAP estimate and a selection-model estimate, and would test whether the slope changes when estimates from instrumental-variable studies are separated from those of ordinary least squares. If all methods place the corrected estimate between -0.5 and -0.9, the conclusion that the raw mean overstates the elasticity is strong; if they give widely different answers, the review reports a range and explains the reasons.
Specialties and sub-fields
Sub-fields differ in their parameters and designs. Pages for sub-fields are added as they are completed.
Health economics
Cost-effectiveness inputs, willingness to pay, and the value of a statistical life.
Development economics
Randomized evaluations, cash transfers, microfinance and aid, with heterogeneous effects.
Environmental economics
Valuation estimates, carbon prices and policy responses.
Labor economics
Returns to schooling, minimum wage and employment elasticities.
Causal designs, structural estimates and comparability
Economic evidence comes from designs that identify effects in different ways. Randomized evaluations, common in development economics, give estimates of average effects for the sampled population. Natural experiments and regression discontinuity designs estimate local effects, for the people whose behavior changed at the threshold or because of the instrument. Panel regressions with fixed effects control for stable differences but not for time-varying confounders. Structural models give parameters that depend on the assumed model. These estimates answer different questions, and averaging them without comment can mislead.
A review should therefore code the identification strategy and test whether it moderates the estimate. It should also pay attention to external validity: a randomized program in one region may yield an effect that differs in another, and the spread of estimates across sites is informative in its own right. Reviews of such programs often model effects with a random-effects structure and report the share of variance between sites. Meta-analyses of cash transfer or microfinance trials have used this approach and have found considerable variation across settings.
Research on the same question by the same authors using the same data should be handled with care. Multiple papers from one data set are not independent, and the review should record data sources so that duplicate samples can be found and handled.
Data collection and replication
Collecting estimates from economics papers takes more work than collecting trial results. Estimates are in regression tables, often with several specifications per table, and the information needed to compute the standard error or the degrees of freedom may be in notes. Coding variables describing the specification is demanding, since papers use different terms. Two coders should work on the data, with a third person checking a sample, and the coding rules should be written before the work starts.
The data and code should be made available. MAER-Net guidelines stress replication: another researcher should be able to start from the posted data and obtain the results in the paper. A good review also states what it does with studies that give too little information to compute an effect size, and tests whether excluding them changes results.
How we support research projects in this area
From estimates to a published meta-regression
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A question, the effect size and conversion rules, the plan for specification variables and a protocol in line with MAER-Net guidelines.
Searching and extraction
Searches of EconLit, Scopus, RePEc and working paper series, with coding of estimates and specifications by two people.
Synthesis
Meta-regression with clustered or multilevel models, FAT-PET-PEESE and model-averaging sensitivity analyses.
Manuscript and submission
The manuscript, replication data and code, and journal preparation.
Boundaries of this service
Meta-analysis in economics summarizes estimates from published studies. Pooled values describe the literature, not the true value in a given country or time, and corrections for selection bias depend on assumptions. We do not give forecasts or investment, policy or legal advice, and a review does not replace the evaluation of a specific policy in its own context.
When the estimates are too different in definition to compare, a structured literature review without pooling may be the better choice, and we say so at the start.
Frequently asked questions
What is meta-regression analysis in economics?
A method that regresses reported estimates on their precision and on characteristics of the studies, to test for publication selection and to explain differences between estimates.
What is FAT-PET-PEESE?
A set of tests and estimators based on regressing estimates on their standard error or variance, used to detect funnel asymmetry and to estimate a selection-corrected effect.
How do I deal with many estimates per study?
Cluster standard errors by study, or use multilevel models, and test several weighting schemes.
Can I use partial correlations?
Yes, when units differ across studies. They can be computed from t-statistics and degrees of freedom, and results can be converted back to an economic scale with the conversion stated.
Where should I search?
EconLit, Scopus, Web of Science and RePEc, with working paper series such as SSRN and NBER.
Do you provide forecasting or policy advice?
No. The service covers research and evidence-synthesis support only.
References
- Stanley TD, Jarrell SB. Meta-regression analysis: a quantitative method of literature surveys. J Econ Surv. 1989;3(2):161-170.
- Stanley TD, Doucouliagos H. Meta-regression approximations to reduce publication selection bias. Res Synth Methods. 2014;5(1):60-78.
- Stanley TD, Doucouliagos H. Meta-regression analysis in economics and business. London: Routledge; 2012.
- Stanley TD, Doucouliagos H, Giles M, et al. Meta-analysis of economics research reporting guidelines. J Econ Surv. 2013;27(2):390-394.
- Havranek T, Stanley TD, Doucouliagos H, et al. Reporting guidelines for meta-analysis in economics. J Econ Surv. 2020;34(3):469-475.
- Ioannidis JPA, Stanley TD, Doucouliagos H. The power of bias in economics research. Econ J. 2017;127(605):F236-F265.