Evidence synthesis in management
Meta-analysis is widely used in organizational research. Studies of job satisfaction and performance, of transformational leadership, of team diversity and of corporate social responsibility and firm results have all been summarized with pooled correlations. A second tradition, systematic review for evidence-informed management, was proposed by Tranfield, Denyer and Smart in 2003, who adapted the medical model to the needs of management research.
The field presents features that matter for reviews. Constructs such as engagement, trust or innovation are measured with different scales and defined in different ways by different research groups. Samples come from firms, teams or individuals in particular industries and countries. Many studies measure all variables at one time through the same questionnaire, which raises concern about common-method bias. Effects are usually small to moderate correlations, and the average effect is less informative than how it differs between contexts.
Our work in this area follows the framework in meta-analysis and systematic review, with the extra care that these features need.
Two traditions of meta-analysis
| Approach | Main idea | Points to check |
|---|---|---|
| Hunter-Schmidt (psychometric) | Correct correlations for unreliability and range restriction, then pool | Needs reliability data; artifact distributions are often assumed; corrected values have larger uncertainty |
| Hedges-Olkin | Pool Fisher-transformed correlations by inverse-variance weights, random effects | Reports uncorrected effects; widely used in medicine and increasingly in management |
| Meta-analytic SEM | Pool a correlation matrix, then fit a structural model to it | Needs enough studies for each pair of variables; sample-size definition is debated |
| Systematic review without pooling | Structured search and thematic synthesis | Suited to heterogeneous constructs and qualitative or case evidence |
Reviews should state which approach they use and why. Corrections for measurement error can change results considerably, so presenting both the uncorrected and the corrected estimate is good practice. Meta-analytic structural equation modeling lets reviewers test a theoretical model built from pooled correlations, but the result depends on how the sample size is defined and on the number of studies that report each correlation.
Judgment calls and replicability
Aguinis and colleagues documented how many choices a meta-analyst makes: which studies to include, how to treat dependent samples, what to do with outliers, how to code moderators and which corrections to apply. Different choices can lead to different conclusions. They argue that these judgment calls should be reported so that others can replicate the review, and that is now an expectation in leading journals.
In practice this means a written protocol, a coding manual with decision rules, a table of every study with the effect size and its source, and a sensitivity analysis showing how results change under the main alternatives. Combs, Crook and Rauch discuss contemporary approaches and unresolved controversies, and we use this as a checklist when planning projects.
Context, moderators and effect-size benchmarks
The average effect in management is often small, and the interesting question is when it is larger or smaller. Common moderators include industry, country or culture, firm size, level of analysis (individual, team or firm), the measurement source (self-report or independent), the research design and the year. Meta-regression handles several moderators together, with the usual limits on power and confounding between moderators.
Bosco and colleagues derived effect-size benchmarks from the distribution of correlations in applied psychology and management, which show that the typical correlation is smaller than Cohen's benchmark of 0.3 for a medium effect would imply. Benchmarks from the specific field give a better frame for judging an effect than general ones. Even then, the practical importance of an effect depends on the cost and the setting, and a small correlation applied to many employees can have a large aggregate effect.
Causal claims need caution. A positive correlation between a firm practice and performance may reflect selection (successful firms adopt the practice) or reverse causation. Reviews should identify which studies use designs that address this, such as longitudinal data, instruments or experiments, and compare them with cross-sectional studies.
Publication bias and common-method bias
Publication bias is a concern in this literature, as in others, and the same tools apply: funnel plots, regression tests, trim-and-fill and selection models. Dissertations and conference papers are a way to find studies that did not reach journals, and a comparison of published and unpublished effects is informative. Common-method bias, caused by measuring predictor and outcome in the same questionnaire from the same respondent, can inflate correlations. Reviews can test it by comparing effects from single-source and multi-source studies.
Research in management has its own history of bias. Studies with hypotheses that fit a theory have been more likely to be published, and the number of significance tests that can be run on survey data is large. Reporting how many of the included effects came from the main hypothesis and how many from exploratory analyses, where this can be coded, helps readers judge the risk.
Search and reporting
Searches use Business Source Complete, ABI/INFORM, Web of Science, Scopus, PsycINFO and EconLit as relevant, with working-paper repositories such as SSRN and dissertation databases. The search often needs many synonyms for each construct, since terminology is not standardized, and a documented strategy is essential. Reporting follows PRISMA 2020 with the extra detail that journals in the field expect on coding and corrections.
Specialties and sub-fields
Sub-fields share methods but differ in their constructs and data. Pages for sub-fields are added as they are completed.
Human resource management
Selection, training, compensation and engagement, with employee and firm outcomes.
Marketing
Consumer behavior, advertising effects and brand outcomes across studies.
Organizational behavior
Attitudes, teams, leadership and motivation in work settings.
Entrepreneurship
Founder characteristics, new venture performance and context.
Supply chain management
Relationships, integration and performance across firms.
Finance & accounting
Event studies, disclosure and market responses, with estimates from archival data.
Strategic management
Strategy, resources, diversification and firm performance.
A worked reading of a pooled correlation
Suppose a review pools correlations between employee engagement and job performance across studies. The figures here are invented to show how a result is read and are not from a real review. The uncorrected random-effects mean correlation is 0.25, with a 95 percent confidence interval of 0.20 to 0.30. After correction for unreliability in both measures the mean is 0.32, with a wider interval. The prediction interval for the uncorrected correlation runs from 0.02 to 0.45.
A careful reader notes four things. The average is positive and the confidence interval excludes zero. The prediction interval shows that in some settings the link could be near zero and in others fairly strong. The corrected value is larger, but it rests on assumed reliabilities. And the studies are mostly cross-sectional, so the figures describe association and do not show that engagement raises performance. The review would then report the moderators, such as whether performance was rated by a supervisor or taken from records, and whether the correlation differs between single-source and multi-source studies.
What a coding manual contains
Coding is the step where management reviews most often differ from each other. A manual written before coding begins lists each variable, its definition, its allowed values and examples. It states how to code a construct that is named differently across papers, for example when "organizational commitment" and "employee loyalty" overlap. It states what to do when a paper reports several samples, several time points or several measures, and how to record the source of every number so that it can be checked.
Two coders work independently on at least a sample of studies, and the agreement is reported (for example with Cohen's kappa for categorical codes and the intraclass correlation for continuous ones). Disagreements are settled by discussion and the manual is revised if a rule was unclear. The final dataset, with the source page for each effect size, is the audit trail for the review. Keeping it in a shared file, and sharing it with the paper where permitted, supports replication.
Moderators coded from the text, such as industry or country, should be recorded with the same care as effect sizes. Missing moderator information is common, and the review states how many studies could not be coded. Sensitivity analyses can show how the results change when studies with uncertain codes are removed.
Using findings in practice
Evidence-informed management asks practitioners to use the best available evidence along with their own knowledge of the organization. A meta-analysis supports this when it reports effects in forms that managers can use: the size of the average effect in terms of an outcome they know, the conditions under which the effect is larger or smaller, and the strength of the evidence. A summary of this kind is more helpful than a statement that a relationship is significant. The review should say plainly where evidence is thin, such as in small firms or in particular regions, so that readers do not apply results where the studies do not reach.
How we support research projects in this area
From a research question to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Protocol and coding manual
A question, inclusion rules, construct definitions and decision rules for judgment calls.
Searching and coding
Searches across business and social-science databases and working-paper repositories, with double coding.
Analysis
Psychometric or inverse-variance pooling, moderator analysis, sensitivity analysis and meta-analytic structural models.
Manuscript and submission
The manuscript, tables of coded studies and journal preparation.
Boundaries of this service
A meta-analysis of management studies summarizes published associations. Most of the evidence is correlational, so pooled effects do not prove that a practice causes performance. Results from one industry or country may not apply in another. We do not provide consulting advice for a particular firm, and a review does not replace the judgment of those who manage an organization.
When the constructs are too varied to pool, we recommend a systematic review or a mapping study instead, and say so at the start.
Frequently asked questions
Should I use Hunter-Schmidt or Hedges-Olkin methods?
Either can be justified. Hunter-Schmidt corrects for measurement artifacts but needs reliability data. Hedges-Olkin is simpler. Many reviews report both.
Can I pool different measures of the same construct?
Only if the construct is defined consistently. The review should code the measures, justify the grouping and test whether the measure type moderates the effect.
How do I handle dependent samples?
Identify studies that report the same sample more than once, combine or select effects by a rule set in advance, and use multilevel or robust variance models for dependent effects.
What is meta-analytic structural equation modeling?
A two-stage method that pools correlation matrices across studies and fits a path model to the pooled matrix.
Do cross-sectional correlations show causal effects?
No. Reviews should separate designs that address causality from those that do not and interpret accordingly.
Do you offer consulting for individual firms?
No. The service covers research and evidence-synthesis support only.
References
- Hunter JE, Schmidt FL. Methods of meta-analysis: correcting error and bias in research findings. 3rd ed. Thousand Oaks (CA): Sage; 2015.
- Tranfield D, Denyer D, Smart P. Towards a methodology for developing evidence-informed management knowledge by means of systematic review. Br J Manag. 2003;14(3):207-222.
- Aguinis H, Dalton DR, Bosco FA, Pierce CA, Dalton CM. Meta-analytic choices and judgment calls: implications for theory building and testing, obtaining replicable results, and cumulative knowledge. J Manag. 2011;37(1):5-38.
- Combs JG, Crook TR, Rauch A. Meta-analytic research in management: contemporary approaches, unresolved controversies, and rising standards. J Manag Stud. 2019;56(1):1-18.
- Bosco FA, Aguinis H, Singh K, Field JG, Pierce CA. Correlational effect size benchmarks. J Appl Psychol. 2015;100(2):431-449.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.