Meta-analysis in sports and exercise science

Sport and exercise research studies training, nutrition, recovery and performance, often in small groups of athletes or volunteers. The trials are short and the samples small, so single studies are imprecise. Meta-analysis helps, if it treats the design of these studies with care.

Evidence synthesis in sport and exercise

A typical sport science study has 10 to 30 participants, lasts 4 to 12 weeks and measures performance with tests such as strength, sprint time, endurance or jump height. Single studies of this kind seldom settle a question. Meta-analysis combines them to estimate average effects and to see how the effect of training changes with its variables: intensity, volume, frequency, participant age and training status.

The literature has some well-known weaknesses. Studies are small and may be underpowered, so the published effects are noisy and likely to be inflated. Blinding is usually impossible. Outcomes are often measured on tests with learning effects, and variation between days is considerable. Many studies report multiple outcomes, and trials in elite athletes are scarce. Caldwell and colleagues called in 2020 for more transparent research practices, including preregistration and data sharing, and the same reasoning applies to reviews.

Our approach follows the framework in meta-analysis and systematic review, with specific attention to the issues below. The medical side of exercise is also covered under health sciences, including rehabilitation.

Effect sizes for small samples and repeated measures

The standardized mean difference is the usual effect size. In small samples it is biased upward, and Hedges' correction, a factor that depends on the degrees of freedom, should be applied. For a single group measured before and after training, the effect can be expressed as the mean change divided by the pre-test standard deviation. For example, if a group of 14 athletes improves from a mean of 52 to 58 units with a pre-test standard deviation of 8, the standardized change is (58 - 52)/8 = 0.75. The small-sample correction is 1 - 3/(4 x 13 - 1) = 0.941, so the corrected value is 0.71. The data are invented.

Morris and DeShon showed how to combine effect sizes from repeated-measures and independent-groups designs on a common metric. The variance of the pre-post effect depends on the correlation between pre and post scores, which is often not reported. Reviews therefore make an assumption (for example, 0.5 or a value taken from test-retest reliability studies), vary it in a sensitivity analysis and report what difference it makes. Crossover trials need a similar treatment, since a paired analysis uses the within-person correlation, and treating them as parallel groups ignores it.

Hopkins and colleagues proposed thresholds for effect sizes in sports and an approach based on magnitude-based inference, which has been criticized for its statistical properties. We use conventional confidence intervals, prediction intervals and, when suitable, Bayesian credible intervals, and we interpret effects against the smallest worthwhile change in the performance measure where one has been defined.

Small studies, bias and the unit of analysis

With small trials, small-study effects are expected: smaller studies show larger effects, for reasons that include publication bias, lower quality and real differences in the intervention. Funnel plots, regression tests and selection models are used as in other fields, with the caution that the number of studies is often small. A statement that bias is absent is not justified by a non-significant test.

The unit of analysis requires attention. Studies often compare several training protocols against one control, several outcomes in the same participants and several time points. Correlated effects should be handled with multilevel models or robust variance estimation. Split-body designs and crossover designs should not be entered as independent groups. For multi-arm trials the control group should be shared appropriately, not counted several times.

Randomization and allocation concealment are often not described. Risk of bias is assessed with RoB 2 for randomized trials and ROBINS-I for non-randomized studies, with an additional note on reporting of training adherence, since an intervention that was not delivered cannot show an effect.

Training variables and dose-response

The practical question for coaches and researchers is not only whether training works but how much, how often and for whom. Meta-regression can relate the effect to training frequency, weekly volume, intensity, program length, age, sex, training status and the type of test. The usual cautions hold: moderators are correlated, study-level data cannot show what happens within individuals, and the number of studies limits the number of moderators.

Dose-response relations are often non-linear. More training can increase gains up to a point and add fatigue or injury risk beyond it. Methods in dose-response meta-analysis, including splines, can describe this. Reviews should present the range of doses in the data and avoid extending curves outside it.

Training status deserves care. Untrained people improve quickly with almost any program, whereas trained athletes improve slowly, so average effects across both groups mix different processes. Stratifying by training status is usual.

Transparency and reporting

Reporting follows PRISMA 2020. Protocols can be registered with the Open Science Framework, and PROSPERO accepts reviews with a health-related outcome. Searches use PubMed, SPORTDiscus, Web of Science, Scopus, Embase and CINAHL, with checks of reference lists and conference abstracts. Records should be kept of all effect sizes and the assumptions used to compute them, and the data and analysis code should be shared where possible. Extraction from figures is common, and the digitizing method and checks between extractors should be reported.

Specialties and sub-fields

Sub-fields differ in measures and populations. Pages for sub-fields are added as they are completed.

Measurement error and the smallest worthwhile change

Performance tests have error. A sprint time or a jump height varies from day to day even when nothing changes, and the size of that variation sets a limit on what can be detected. Test-retest reliability, usually reported as a coefficient of variation or a typical error, is therefore useful information for interpreting results. If the average gain from a training program is smaller than the typical error of the test, a single athlete cannot tell whether their own change reflects training or noise, although the group average may still be estimated accurately.

The smallest worthwhile change is the smallest improvement that matters for performance. In sports where results are decided by fractions of a percent, a change of 0.3 to 0.5 percent may be worthwhile, while in general fitness research a much larger change may be needed. Where authors define such a threshold, the review can express effects in relation to it. Without a threshold, the standardized effect is hard to interpret, and the review should say how it is read and why.

Different tests of the same quality, such as several measures of strength, can be combined only if they are close in meaning. A one-repetition maximum, an isokinetic torque and a jump are not interchangeable. Reviews usually analyze them by outcome type and test whether the type of test moderates the effect.

Supplements, nutrition and placebo effects

Studies of supplements, ergogenic aids and diets are common. Because participants often know what they are taking, expectations influence results, and effects reported in open-label studies tend to be larger than those in placebo-controlled trials. Reviews should separate the two or test the design as a moderator. Studies funded by manufacturers may report larger benefits, and funding should be recorded.

Dose, timing and baseline status matter. Caffeine, creatine and nitrate studies, for example, differ in the amounts used and in the habitual intake of the participants. A response that depends on the baseline intake cannot be summarized by a single mean without noting this. For safety outcomes, the review should look for adverse events and for any contamination of supplements, which is a recognized issue for athletes subject to anti-doping rules. Such rules and product safety are outside the scope of a synthesis of effects, and athletes should consult their governing body and a qualified professional.

Injury and observational evidence

Injury prevention is studied in randomized trials of exercise programs and in cohort studies of injury risk. Trials of team-based prevention programs are often cluster-randomized by team, and the review should use effect sizes adjusted for clustering. Outcomes such as injury incidence per 1,000 hours of exposure are rates, and they should be pooled as rate ratios using exposure time and not as proportions of players. Observational cohorts report associations between training load, previous injury or movement screening and later injury; such risk factors are often weak predictors, and pooled estimates are subject to confounding and to the way injury was defined. Consensus statements on injury definitions help to harmonize studies, and the review should state which definition each used.

For athletes with a health condition, or for rehabilitation, the evidence is covered under health sciences and its specialties.

Elite athletes and generalization

Most trials use students, recreational athletes or volunteers, because elite athletes are few, busy and reluctant to alter their training. Reviews should state the populations studied and avoid applying the average to elite athletes unless the data include them. Effects usually shrink with training level: a program that improves strength greatly in beginners may add little in people who have trained for years. Studies of female athletes and of older and younger athletes are less common, and the gaps should be reported. Where the review aims to inform practice, it can add a short section on how the findings might be applied, with a note that individual response varies and that testing a program with the athlete's own monitoring is more informative than relying on the average from a research study.

How we support research projects in this area

Support

From small trials to a published synthesis

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A question, effect sizes for repeated-measures designs, assumptions about correlations and a protocol for registration.

  • Searching and extraction

    Searches of sport and health databases, with extraction from figures and checks between extractors.

  • Synthesis

    Random-effects and multilevel models, dose-response and meta-regression, and small-study analyses.

  • Manuscript and submission

    The manuscript, data and code, and journal preparation.

Get a quoteDescribe the intervention, outcomes and target journal.

Boundaries of this service

A meta-analysis summarizes studies of groups, mostly small ones. A pooled effect does not tell a coach or an athlete what will work for one person, and training response varies between people. We do not provide individual training, medical or rehabilitation advice. The service supports research and does not replace the judgment of qualified coaches and clinicians.

If studies are too varied to pool, we recommend a systematic review without meta-analysis and say so at the start.

Frequently asked questions

Which effect size suits training studies?

Hedges' g for between-group differences, and a standardized change score for pre-post designs, with the assumed pre-post correlation tested in a sensitivity analysis.

Why correct for small samples?

The standardized mean difference is biased upward in small samples, and Hedges' correction reduces this.

How do I handle crossover trials?

Use a paired analysis with the within-person correlation, or the effect size from the paired comparison, and not an analysis that treats the periods as separate groups.

Can I find the dose that works best?

Dose-response models can describe the pattern within the range of doses studied. They cannot show the best dose for an individual.

What is the minimum number of studies?

There is no fixed number. With few studies, moderator analysis and bias tests have little power, and the review should say so.

Do you give training or medical advice?

No. The service covers research and evidence-synthesis support only.

References

  1. Hopkins WG, Marshall SW, Batterham AM, Hanin J. Progressive statistics for studies in sports medicine and exercise science. Med Sci Sports Exerc. 2009;41(1):3-13.
  2. Morris SB, DeShon RP. Combining effect size estimates in meta-analysis with repeated measures and independent-groups designs. Psychol Methods. 2002;7(1):105-125.
  3. Lakens D. Calculating and reporting effect sizes to facilitate cumulative science: a practical primer for t-tests and ANOVAs. Front Psychol. 2013;4:863.
  4. Caldwell AR, Vigotsky AD, Tenan MS, et al. Moving sport and exercise science forward: a call for the adoption of more transparent research practices. Sports Med. 2020;50:449-459.
  5. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.