Meta-analysis and evidence synthesis for finance and accounting research

Finance and accounting research studies markets, firm decisions, reporting and governance, mostly with archival data. The evidence is observational, and the same databases are used again and again, so results are not independent in the way trials are. Reviews have to deal with model dependence, multiple testing, overlapping samples and the difference between statistical and economic significance.

Evidence synthesis in finance and accounting

Finance and accounting research asks questions such as whether board independence relates to firm value, how markets react to earnings announcements or mergers, whether stronger disclosure lowers the cost of capital, and how auditor characteristics relate to reporting quality. Synthesis in this field is less common than in medicine or psychology, but meta-analyses exist on topics such as corporate social responsibility and financial performance, governance and performance, and merger returns. Many reviews in the area are narrative or bibliometric.

The evidence has features that call for care. Almost all studies use observational archival data, usually drawn from a few common databases, so samples overlap heavily. Estimates come from regression models with different controls, and many tests are run per paper. Publication favors significant results. Our methods follow meta-analysis and systematic review practice, adapted to these features and informed by guidance from the Meta-Analysis of Economics Research Network. This page builds on the general guidance for management and business.

Effect measures from regressions

Common effect measures
MeasureTypical useIssue for synthesis
Regression coefficientEffect of a variable on an outcome, in original unitsUnits and functional form differ; coefficients cannot be pooled across specifications
Partial correlationStandardized effect derived from a coefficient and its t-statisticComparable across studies, but requires t-statistics and degrees of freedom to be reported
Abnormal returnExcess return around an event, relative to a benchmark modelDepends on event window, benchmark and sample; cumulative returns grow with window length
ElasticityPercentage change in outcome per percentage change in predictorComputed at means; sensitive to variable construction
t-statisticReported significance of a coefficientUsed in p-curve and funnel-asymmetry tests; not an effect size

Regression coefficients from different models cannot be pooled because they are in different units. The partial correlation, computed from a t-statistic and the residual degrees of freedom, puts them on a common scale: r = t divided by the square root of t squared plus the degrees of freedom. It is widely used in meta-analyses of economics and finance research. Its drawback is that it describes the strength of association after other controls, so the pooled value depends on which controls studies included, which should be coded and tested.

Event studies

Event studies measure the market reaction to an announcement by comparing observed returns with expected returns from a benchmark model such as a market model. The result is an abnormal return, often cumulated over a window of a few days around the announcement. Meta-analysis of event studies pools these abnormal returns, but their comparability depends on the event window, the benchmark model, the market, the sample period and whether the event was anticipated. A review should code each, test them as moderators and be cautious about pooling across very different events.

Events cluster in time (for example, merger waves) and across firms in the same industry, which creates cross-sectional dependence. Standard errors that ignore this are too small. A review should note how each study handled it and should avoid giving large weight to studies with understated standard errors.

Overlapping samples and common databases

Papers in finance and accounting use the same databases of stock returns, firm accounts, analyst forecasts and governance indicators, often for the same countries and periods. Two papers asking similar questions can analyze nearly the same observations, so their estimates are correlated. Counting them as independent understates uncertainty. A review should record the database, country and period for each estimate, and use methods such as robust variance estimation or multilevel models with sample clusters. Sensitivity analyses that drop overlapping studies indicate how much the result depends on this.

The same concern applies to a single paper that reports dozens of specifications. Selecting one by a rule stated in advance, or modelling the dependence, is better than treating all as independent.

Publication bias, p-hacking and replication

Studies in economics and finance have been found to contain more significant results than expected, and tests for funnel-plot asymmetry, p-curves and selection models are standard tools for assessing this. The precision-effect test and related methods estimate the effect that would remain at infinite precision, and their results can be much smaller than the raw average. These tests make assumptions and can mislead when heterogeneity is large, so we report several and explain their limits.

Replication in finance has shown that some of the many return predictors reported in the literature fail out of sample or after publication, or when multiple testing is accounted for. A review of a topic with many tested predictors should consider multiple-testing corrections and should compare effects before and after publication where data allow.

Endogeneity and causal claims

Governance, disclosure and capital structure are chosen by firms, so correlations with performance are open to reverse causation and omitted variables. Better firms may adopt better governance. Studies using natural experiments, instrumental variables, regression discontinuity or difference-in-differences make stronger causal claims, but their assumptions can be challenged. A review should code identification strategy and compare estimates across strategies. A typical result is that associations from simple regressions are larger than those from designs that address endogeneity. We say so plainly and avoid causal wording for correlational estimates.

Corporate social responsibility and governance as examples

Meta-analyses on corporate social responsibility (CSR) and financial performance have found small positive average associations, with wide variation. Measures of CSR come from different rating agencies with low agreement among them, which means that results depend on the data vendor. Financial performance is measured by accounting returns, market returns or value measures. A review should treat each combination separately and report how much results change by measure. Governance studies show similar variation, with mixed results for board independence, CEO duality and ownership structure. The point of a synthesis is to display the pattern and its dependence on measures, not to produce a single verdict.

Experimental and behavioral accounting research

Accounting also includes laboratory and survey experiments with students, auditors and managers, on topics such as judgment, ethics and the effect of disclosure on decisions. These produce standardized mean differences and can be pooled as in psychology, with moderators for participant type (students versus professionals), task and incentive. Effects from students may differ from those of experienced professionals, and the review should test this. Studies using hypothetical scenarios differ from those with real stakes, and this should be coded.

A worked conversion

To show how a partial correlation is obtained, take an invented regression with a t-statistic of 2.5 and 200 residual degrees of freedom. The partial correlation is 2.5 divided by the square root of 6.25 plus 200, which is about 0.174. A second invented study with t of 1.2 and 4,000 degrees of freedom gives 1.2 divided by the square root of 1.44 plus 4,000, about 0.019. The large study, although less significant, is more precise about a small effect, and a weighted analysis would reflect this.

This shows why the procedure is useful: both studies are now on one scale, and the review can ask whether the difference between 0.17 and 0.02, which are the two values after rounding, is explained by sample size, data, controls or chance. It also shows its dependence on reporting. If a paper omits t-statistics or degrees of freedom, the effect cannot be computed, and the reviewer must request them from authors or exclude the study with a stated reason that is recorded in the flow diagram.

Common pitfalls we look for

  • Pooling raw coefficients from different units and models.
  • Treating overlapping database samples as independent.
  • Ignoring event window and benchmark choice in event studies.
  • Counting every specification in a paper as a separate study.
  • Interpreting statistical significance as economic importance.
  • Causal language for associations without an identification strategy.
  • Not testing for p-hacking in topics with many tested predictors.

Planning a finance or accounting synthesis

We help define the relationship being studied, choose between partial correlations, abnormal returns and other measures, plan searches of finance, accounting and economics databases and working-paper repositories (SSRN, NBER, RePEc), and set up coding of database, country, period, controls and identification strategy. We register the protocol and follow reporting guidance of the economics meta-analysis community together with PRISMA 2020 where it applies. See the meta-analysis service for scope and process.

Statistical and economic significance

Suppose a review reports a pooled partial correlation of 0.05 between a governance index and firm value, with a prediction interval from minus 0.04 to 0.14. These numbers are invented. A correlation of 0.05 is small by conventional benchmarks, and with large samples it can be highly significant. The prediction interval includes zero, so in some settings the true association may be negative or nil. Whether the average is economically meaningful depends on the scale of the variables: a unit change in a governance index might correspond to a small or a large change in the firm's governance. The review should translate the effect into units that readers in the field would recognize, such as a change in firm value for a given change in the index, and state the uncertainty.

How we support research projects in this area

Support

From a research question to a published synthesis

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Protocol and coding manual

    A question, inclusion rules, construct definitions and decision rules for judgment calls.

  • Searching and coding

    Searches across business and social-science databases and working-paper repositories, with double coding.

  • Analysis

    Psychometric or inverse-variance pooling, moderator analysis, sensitivity analysis and meta-analytic structural models.

  • Manuscript and submission

    The manuscript, tables of coded studies and journal preparation.

Get a quoteDescribe your constructs, data and target journal.

Boundaries of this service

A synthesis of finance and accounting research describes average associations across studies and databases. It does not provide investment advice, trading signals, audit opinions or valuation of any firm, and past patterns in the literature do not predict future returns. Most data are observational and overlapping, so causal conclusions require strong designs and the results may not generalize across markets or periods.

Frequently asked questions

Can regression coefficients be meta-analyzed?

Not directly if units differ. Partial correlations computed from t-statistics put them on a common scale, with attention to which controls each model used.

How do you treat event study results?

As abnormal returns, with moderators for window, benchmark and sample, and with attention to clustering of events in time.

Why is overlap of samples a problem?

Papers using the same database and period have correlated estimates, so treating them as independent understates uncertainty.

Do you test for p-hacking?

Yes, with funnel-asymmetry tests, p-curve and selection models, explaining the limits of each.

Does meta-analysis show whether governance causes performance?

Only when included designs address endogeneity. We compare estimates across identification strategies.

Do you provide investment advice?

No. The service provides research and evidence-synthesis support only.

References

  1. Stanley TD, Doucouliagos H. Meta-regression analysis in economics and business. London: Routledge; 2012.
  2. Stanley TD, Doucouliagos H, Giles M, et al. Meta-analysis of economics research reporting guidelines. J Econ Surv. 2013;27(2):390-394.
  3. MacKinlay AC. Event studies in economics and finance. J Econ Lit. 1997;35(1):13-39.
  4. Orlitzky M, Schmidt FL, Rynes SL. Corporate social and financial performance: a meta-analysis. Organ Stud. 2003;24(3):403-441.
  5. Harvey CR, Liu Y, Zhu H. ... and the cross-section of expected returns. Rev Financ Stud. 2016;29(1):5-68.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.