Guide

Meta-analysis of correlation coefficients

Correlations are the common currency of psychology, education, management and many observational fields, and pooling them is straightforward once the right transformation is used. This guide explains Fisher's z, works through an example, and covers the practical issues that matter more than the arithmetic: measurement error, restricted range and dependent effects.

When correlations are the effect size

A correlation coefficient, r, measures the strength and direction of a linear association between two continuous variables, from -1 to 1. A meta-analysis of correlations answers questions of the form: how strongly is job satisfaction related to performance, or reading ability to later mathematics achievement, or a biomarker to a clinical score? The studies each report an r with a sample size, and the synthesis estimates an average association and how much it varies.

Correlations are the usual effect size when there is no intervention or exposure contrast, only an association between two measured variables. Where studies are experimental, a standardized mean difference or a ratio measure is more natural, and the two families can be converted when they must be combined, with attention to the assumptions.

Why Fisher's z

The sampling distribution of r is not symmetric, especially when the true correlation is far from zero, and its variance depends on the correlation itself. A study with r = 0.8 has a smaller sampling variance than one with r = 0.2 for the same sample size, so r cannot be pooled by simple inverse-variance methods without bias.

Fisher's z transformation, z = 0.5 x ln((1 + r) / (1 - r)), which is the inverse hyperbolic tangent of r, makes the distribution approximately normal and gives a variance that depends only on the sample size: 1 / (n - 3). The weight of a study is therefore n - 3. This simplicity is why the transformation is standard.

The steps are: transform each r to z, pool the z values with weights n - 3 under a fixed-effect or random-effects model, find the confidence interval for the pooled z, and back-transform to r, with r = (e2z - 1) / (e2z + 1). All the confidence limits are back-transformed too. The intervals for r are then asymmetric, which is correct.

A worked example

Four simulated studies report these correlations between two measures.

Four simulated studies, with Fisher's z and the standard error of z
StudynrFisher's zSE of zWeight (n - 3)
Study A1000.300.3100.10297
Study B600.450.4850.13257
Study C2000.220.2240.071197
Study D800.380.4000.11477

The weighted mean of z is 0.310, with a standard error of 0.048. Back-transforming gives a pooled correlation of 0.30, with a 95 percent confidence interval of 0.21 to 0.38. Cochran's Q is 3.83 on 3 degrees of freedom. The simple unweighted mean of the four correlations is 0.34, and it gives the small studies as much influence as the large one, which is why weighting matters. The studies are simulated for illustration.

Heterogeneity and model choice

Correlations often vary between studies for substantive reasons: different populations, instruments, settings and ranges of the variables. A random-effects model is usually the more defensible choice, and the report should give tau-squared on the z scale and a prediction interval. The usual cautions on I-squared apply, and with few studies the estimates are imprecise.

Moderator analyses, by subgroup or meta-regression, can explore sources of variation, for example the type of instrument, the age of participants or the country. With correlations there is a particular source of apparent heterogeneity that has nothing to do with the underlying relationship: differences in range and reliability, discussed next.

Range restriction and measurement error

An observed correlation depends on how the variables were measured and on who was sampled. Two issues stand out.

Measurement error. If either variable is measured with error, the observed correlation is lower than the correlation between the true scores. This is called attenuation. Spearman's formula for correction divides the observed correlation by the square root of the product of the two reliabilities. For an observed r of 0.30 with reliabilities of 0.70 and 0.80, the disattenuated value is 0.30 / sqrt(0.70 x 0.80) = 0.40. The correction rests on reliability estimates that are themselves uncertain, and it increases the variance of the estimate.

Range restriction. If a sample is selected on one variable, for example only admitted students or only employed people, the observed correlation is reduced, because the variable has less spread. Corrections exist but require knowledge of the unrestricted standard deviation.

The psychometric tradition of meta-analysis, associated with Hunter and Schmidt, builds corrections for these artifacts into the method. Other traditions pool observed correlations and treat these as sources of heterogeneity to be described. Both are legitimate, and the choice should be stated in the protocol, with the correction assumptions made explicit. A pooled disattenuated correlation answers a different question from a pooled observed correlation: the first concerns the relationship between constructs and the second concerns the relationship between the measures as used.

Dependent correlations

Many primary studies report more than one correlation from the same sample: several measures of the same construct, several time points, or a matrix of intercorrelations. These correlations are not independent, and treating them as separate studies overstates the precision. The options include selecting one correlation per sample by a prespecified rule, averaging within a study, using robust variance estimation, or fitting a multilevel (three-level) model that accounts for the nesting of effect sizes within studies. Multilevel models and robust variance estimation are now common in the social-science literature. The rule for handling dependence belongs in the protocol, and the approach should be described in the report.

A related problem arises when the same data appear in more than one publication. The overlapping reports must be identified and counted once.

Interpreting a pooled correlation

Cohen's benchmarks of 0.1, 0.3 and 0.5 for small, medium and large correlations are quoted widely. As with standardized mean differences, they are conventions, and empirical work suggests that in some fields, such as individual-differences research, typical values are lower. A pooled correlation of 0.30 explains about nine percent of the variance, and that figure can sound unimpressive or impressive depending on the context. Interpretation is best tied to the field and the consequences, for example by comparing with correlations from other well-established relationships in the same area.

A correlation also does not describe the effect of changing one variable, it does not imply causation, and it can be driven by outliers or by a non-linear relationship. When a pooled association is used to support a causal claim, the argument needs to come from the design of the studies, not from the size of r. For the same reason, a meta-analysis of correlations does not by itself tell clinicians or managers what to do.

Common mistakes

  • Averaging r directly without weights or without the z transformation.
  • Using n instead of n minus 3 for the weights, which is a small error with large samples and a larger one with small samples.
  • Forgetting to back-transform the confidence limits, or reporting limits on the z scale as if they were correlations.
  • Counting several correlations from one sample as independent studies.
  • Mixing partial and zero-order correlations or standardized regression coefficients, which are not interchangeable.
  • Correcting for attenuation without saying so, or with reliabilities that do not fit the sample.
  • Interpreting the result causally.

Reporting

Report the number of studies and independent samples, the sample sizes, the transformation, the model, the estimate of tau-squared, the interval for the pooled r and the prediction interval, any corrections, the handling of dependent effects, and moderator analyses. A forest plot of correlations with intervals, on the r scale after back-transformation, helps readers. The reporting of the review follows PRISMA 2020, and for observational data, the MOOSE checklist is also relevant.

Extracting correlations from what studies report

Primary studies do not always report a plain Pearson correlation, and extraction decisions affect the pooled value. Spearman rank correlations can be pooled with Pearson correlations in many reviews, but they are not the same quantity, and a conversion or a sensitivity analysis is wise. Point-biserial and phi coefficients are correlations too, and they are bounded by the proportions in the groups, so they are not comparable with correlations between continuous variables without care. Standardized regression coefficients from multivariable models are not correlations. Treating a beta weight as r mixes an adjusted with an unadjusted association, and in some fields a rough conversion is used only when no alternative exists, with a sensitivity analysis excluding it.

Where only a p value and a sample size are given, the correlation can be recovered, and where only a sign and a statement of significance are given the study provides almost no usable information. Record the sample size actually used in the analysis, since missing data often reduce it below the number recruited. Note the direction of each scale, and reverse the sign where a high score on one instrument means a low level of the construct on another. Two people should extract the correlations independently, because sign errors and transcription errors are among the most common faults in this kind of review.

Few studies and small samples

The normal approximation underlying Fisher's z works well when samples are not tiny. With very small samples, below about ten participants, the approximation is poorer and the n minus 3 weight can become zero or negative, so such studies need special handling or exclusion. With few studies, the between-study variance is poorly estimated, and a random-effects interval may be too narrow. The Hartung-Knapp adjustment, a prediction interval and a clear statement of the uncertainty are more useful than a precise-looking single figure. Publication bias deserves attention too, because correlation studies are often reported selectively when they reach significance, and small studies with large correlations are a common sign. A funnel plot on the z scale, with the standard error of z, is the usual display, and the usual limits on interpretation apply.

How we can help

We can extract and harmonize correlations, apply Fisher's z pooling with fixed-effect or random-effects models, handle dependent effect sizes with multilevel or robust variance methods, correct for artifacts where the protocol calls for it, and report heterogeneity and moderators. [OWNER VERIFICATION REQUIRED] The relevant services are meta-analysis and statistical analysis.

Frequently asked questions

Why not just average the correlations?

Because r has a skewed sampling distribution and a variance that depends on its value. Fisher's z gives an approximately normal quantity with weights of n minus 3.

How do I get the confidence interval for the pooled correlation?

Calculate the interval on the z scale and back-transform both limits to r. The interval is asymmetric on the r scale.

Should I correct for measurement error?

It depends on the question. Correction estimates the relationship between true scores, but it needs reliabilities and increases uncertainty. State the choice in the protocol.

What do I do with several correlations from one study?

Do not treat them as independent. Use a prespecified rule to choose, average them, or model the dependence with a multilevel or robust variance approach.

Are 0.1, 0.3 and 0.5 appropriate benchmarks?

They are conventions. Interpret a correlation in the context of the field and the consequences, and avoid causal readings.

Can I combine correlations with standardized mean differences?

Conversions exist and rely on assumptions such as normality and equal group sizes. Do it only when necessary and run a sensitivity analysis.

References

  1. Borenstein M, Hedges LV, Higgins JPT, Rothstein HR. Introduction to Meta-Analysis. Wiley; 2009.
  2. Fisher RA. On the probable error of a coefficient of correlation deduced from a small sample. Metron. 1921;1:3-32.
  3. Hunter JE, Schmidt FL. Methods of Meta-Analysis: Correcting Error and Bias in Research Findings. 3rd ed. Sage; 2015.
  4. Spearman C. The proof and measurement of association between two things. Am J Psychol. 1904;15(1):72-101.
  5. Cohen J. Statistical Power Analysis for the Behavioral Sciences. 2nd ed. Lawrence Erlbaum Associates; 1988.
  6. Cheung MWL. A guide to conducting a meta-analysis with non-independent effect sizes. Neuropsychol Rev. 2019;29(4):387-396.
  7. Hedges LV, Tipton E, Johnson MC. Robust variance estimation in meta-regression with dependent effect size estimates. Res Synth Methods. 2010;1(1):39-65.
  8. Gignac GE, Szodorai ET. Effect size guidelines for individual differences researchers. Pers Individ Dif. 2016;102:74-78.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.