Meta-analysis in psychology and behavioral sciences

Psychology is one of the main users of meta-analysis, and one of the fields where its limits have been studied most closely. Effects are small, measures vary, and publication bias and flexible analysis are well documented. We support reviews that report these issues openly.

Evidence synthesis in psychology

Psychology adopted meta-analysis early. Smith and Glass pooled psychotherapy outcome studies in 1977, and Hunter and Schmidt developed methods for correcting correlations for measurement error and range restriction. Today meta-analyses of interventions, individual differences, and basic experimental effects appear in every subfield, and they often shape textbooks and policy.

The field has also been a source of concern. The Open Science Collaboration's 2015 attempt to replicate 100 published studies found that many effects were smaller than first reported, and this focused attention on bias in the published record. Meta-analysts must now ask whether the studies they pool are a fair sample of all the studies done. Methods for this have grown quickly, and the lessons apply beyond psychology.

Our approach combines the standard framework described under meta-analysis with the particular needs of behavioral data: many measures of the same construct, repeated measurements, multiple outcomes per study and effects that are small enough that bias matters.

Effect sizes and what they mean

Effect measures in psychological research
MeasureTypical usePoints to check
Standardized mean difference (Hedges' g)Group comparisons on different instrumentsSmall-sample correction; choice of standard deviation; change scores versus post-test scores
Correlation (Fisher z)Associations between traits, attitudes or behaviorsReliability and range restriction; whether the coefficient is zero-order or partial
Odds ratioBinary outcomes such as remission or dropoutConversion to other effect sizes assumes distributions; interpretation for common outcomes
Pre-post changeSingle-group designsNeeds the pre-post correlation, which is rarely reported

Benchmarks such as Cohen's small, medium and large (0.2, 0.5 and 0.8) were suggested for use when nothing else is known. They are widely quoted and often misapplied, since the effects typical of a field may be well below these values. A better practice is to interpret effects against the distribution of effects in similar research and against a meaningful change on a familiar scale. The Hunter-Schmidt approach corrects correlations for unreliability and range restriction; these corrections need information on reliability that is not always available and can add uncertainty, so the corrected and uncorrected results are both usually reported.

Many outcomes, many measures, three levels

A typical psychological study reports several outcome measures for the same participants, several time points and sometimes several experimental conditions. Entering all of them as independent effects overstates precision. Three-level meta-analytic models, with sampling variance at level one, variation within studies at level two and variation between studies at level three, are one answer; Cheung and Assink and Wibbelink give practical guides for fitting them. Robust variance estimation is another, and it is useful when the correlation between outcomes within a study is unknown.

A related choice is which effect to include when a study measures a construct in several ways. Averaging within the study, selecting a primary measure by a rule fixed in advance and modeling all effects with a dependency structure each have uses. The rule should be stated in the protocol, since selecting the largest effect after seeing the data introduces bias.

Publication bias, questionable practices and corrections

In a literature where significant results are favored, and where researchers have freedom in choosing analyses, the average published effect is biased upward. Meta-analysts therefore examine small-study effects with funnel plots and regression tests, apply selection models, and use methods such as p-curve, p-uniform and PET-PEESE. Carter and colleagues compared these methods in a simulation study in 2019 and found that their performance depends on the amount of bias, the heterogeneity and the number of studies, and that no single method works best in every condition.

This advice follows. Report several methods, explain what each assumes and treat the results as a sensitivity analysis. Where the corrected estimate differs a great deal from the uncorrected one, say so and discuss the reasons. Include unpublished studies where possible, using preprints, dissertations, registered reports and requests to authors. Registered reports, which are accepted before results are known, are a useful comparison group where available.

Lakens and colleagues gave six recommendations for the reproducibility of meta-analyses, including sharing the search strategy, the coding and the analysis code. We apply them as a default: a documented search, a codebook, a table of every coded effect with its source, and the analysis code, so that another team can reproduce each number in the paper without contacting the authors.

Moderators, heterogeneity and the problem of small samples

Psychological effects often vary by population, task, measure and setting. Heterogeneity is expected and should be reported with the between-study variance and a prediction interval. Moderator analysis is common, and it has the usual limits: each moderator is a study-level variable, studies differ in many ways at once and subgroups can be small. A frequent weakness is testing many moderators without adjustment and then reporting those that reached significance.

Samples in psychology are often convenience samples from a few countries. A review can record the sample origin and test whether the results differ, but it cannot create diversity that the primary studies lack. This limit should be stated in the discussion.

For interventions, the treatment effect may depend on the control condition. Waitlist, attention placebo and active treatment controls produce different effect sizes, and combining them may hide this. Comparisons should be stratified by control type.

Designs and what each can show

Psychological evidence comes from designs that answer different questions, and a review should keep them distinct. Randomized experiments, in the laboratory or the field, support causal statements about the manipulated factor, with the caveat that laboratory tasks may not resemble everyday behavior. Randomized trials of psychological interventions face special problems: participants and therapists cannot be blinded to the treatment, outcomes are often self-reported, and the control condition varies. Quasi-experimental studies use matching or statistical adjustment and are open to confounding. Correlational surveys describe associations and cannot separate cause from selection.

Longitudinal studies add the order of events. A cross-lagged panel design, for example, tests whether an earlier measure predicts a later one while controlling for the earlier value of the later one, although it still cannot remove unmeasured confounding. When a review includes several designs, it should analyze them separately or test design as a moderator. Risk of bias is assessed with tools matched to the design, such as RoB 2 for randomized trials and ROBINS-I for non-randomized studies of interventions, and with study-level indicators such as preregistration, sample size justification and use of validated measures.

Replication is also a design feature. Where preregistered replications of an effect exist, they give a check on the earlier literature, and reviews can compare their average with that of the original studies. Large multi-site replication projects have shown that effects often shrink, and a review that includes them tells readers more than one that does not.

Search, coding and reporting

Searches usually include PsycINFO, MEDLINE, Web of Science and Scopus, plus databases of dissertations and preprint servers such as PsyArXiv. Two coders extract the data, using a codebook with definitions and examples, and agreement is reported. The APA Meta-Analysis Reporting Standards, together with PRISMA 2020, give a checklist of items to report. Preregistration of the protocol on a platform such as the Open Science Framework is common, and PROSPERO accepts reviews with a health-related outcome.

Specialties and sub-fields

Sub-fields share the same methods but differ in their designs and measures, and each page below explains what is particular to that part of psychology. Pages for sub-fields are added as they are completed.

A worked reading of a pooled correlation

Suppose a review pools correlations between a measure of sleep quality and a measure of attention. After converting each correlation to Fisher's z, weighting by the inverse of the variance and back-transforming, the pooled correlation is reported with a confidence interval and a prediction interval. The numbers below are invented to show how the reading goes and are not from a real review.

Imagine a pooled correlation of 0.20 with a 95 percent confidence interval of 0.14 to 0.26, and a prediction interval of -0.05 to 0.43. The confidence interval describes the precision of the average. The prediction interval describes the range of true correlations in a new setting. It includes zero and a moderate value, so the average is clearly positive, but a new study could find nothing or a fairly strong link. A correction for unreliability might push the average up, and a bias correction might push it down. A good report presents all three and explains the reasoning, instead of choosing one number to headline.

How we support research projects in this area

Support

From a research question to a published synthesis

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Protocol and preregistration

    A structured question, a coding manual, a plan for dependent effects and a protocol prepared for preregistration.

  • Searching and coding

    Searches across psychological databases and preprint servers, with double coding and an agreement check.

  • Analysis

    Random-effects and three-level models, moderator analysis and a set of publication-bias sensitivity analyses.

  • Manuscript and submission

    Reporting checklists, the manuscript and shareable data and code.

Get a quoteDescribe your constructs, designs and target journal.

Boundaries of this service

Meta-analysis summarizes what published studies found. In psychology, that record is affected by publication bias and by flexible analysis, and statistical corrections are approximate. A pooled effect does not show that an intervention will work for an individual, and we do not provide diagnosis, therapy or clinical advice. The service supports research and does not replace professional judgment.

Where a question is better served by a systematic review without pooling, because measures are too different to combine, we say so at the start.

Frequently asked questions

Which effect size should I use for psychological data?

Hedges' g for group differences, Fisher z for correlations, and odds ratios for binary outcomes. The choice follows the question and should be fixed in the protocol.

How do I handle several outcomes per study?

Use a three-level model or robust variance estimation, or select one effect per study by a rule set in advance.

What are the best methods for publication bias?

No single method is best in all conditions. Use funnel-plot methods, selection models and p-value based methods as a set of sensitivity analyses and report them together.

Should I correct correlations for reliability?

It can be informative if reliability data are available, but the corrections add uncertainty. Report corrected and uncorrected results.

Can you include unpublished studies?

Yes. Searches can include dissertations, preprints and registries, and authors can be contacted for data, subject to the scope agreed.

Do you give therapeutic or diagnostic advice?

No. The service provides research and evidence-synthesis support only.

References

  1. Smith ML, Glass GV. Meta-analysis of psychotherapy outcome studies. Am Psychol. 1977;32(9):752-760.
  2. Hunter JE, Schmidt FL. Methods of meta-analysis: correcting error and bias in research findings. 3rd ed. Thousand Oaks (CA): Sage; 2015.
  3. Open Science Collaboration. Estimating the reproducibility of psychological science. Science. 2015;349(6251):aac4716.
  4. Carter EC, Schonbrodt FD, Gervais WM, Hilgard J. Correcting for bias in psychology: a comparison of meta-analytic methods. Adv Methods Pract Psychol Sci. 2019;2(2):115-144.
  5. Lakens D, Hilgard J, Staaks J. On the reproducibility of meta-analyses: six practical recommendations. BMC Psychol. 2016;4:24.
  6. Cheung MWL. A guide to conducting a meta-analysis with non-independent effect sizes. Neuropsychol Rev. 2019;29(4):387-396.
  7. Assink M, Wibbelink CJM. Fitting three-level meta-analytic models in R: a step-by-step tutorial. Quant Methods Psychol. 2016;12(3):154-174.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.