Evidence synthesis in exercise science
Exercise science examines how the body responds and adapts to exercise and how physical activity relates to health. Topics include resistance training and muscle strength, endurance and high-intensity interval training, flexibility, recovery strategies, nutrition and supplements in sport, physical activity and chronic disease, and exercise in older adults and clinical populations. Meta-analyses are frequent, partly because individual studies are small and results are inconsistent.
The field has distinctive features. Sample sizes are often 10 to 30 participants. Many studies are crossover or pre-post designs. A single paper reports many outcomes measured at several time points. Exercise dose can be described by frequency, intensity, duration, volume and type, each measured in different ways. Our methods follow meta-analysis and systematic review practice, adapted to these features. This page builds on the general guidance for sports and exercise sciences.
Designs and how effects are computed
| Design | Effect size source | Issue for synthesis |
|---|---|---|
| Parallel-group trial | Difference in change (or post) between groups | Small samples; baseline differences; use of change scores or ANCOVA |
| Crossover trial | Within-person difference between conditions | Needs correlation between conditions; carry-over and washout; avoid treating as parallel |
| Single-group pre-post | Change from baseline | No control; improvement may reflect learning or season; not comparable to controlled effects |
| Cohort study of activity | Hazard or risk ratio by activity level | Self-reported activity; confounding; dose from questionnaires |
| Field study of athletes | Differences in performance measures | Sample selection; season timing; unequal training exposure |
Crossover trials need special handling. The standard error of the difference depends on the within-person correlation between conditions, which is rarely reported, so reviews often assume a value and test sensitivity to it. Treating a crossover trial as parallel-group data ignores the pairing and understates precision, while treating the two conditions as separate groups counts participants twice. Pre-post effects within one group should be analyzed apart from controlled comparisons. Standardizing by the pre-test standard deviation or by the change score standard deviation gives different values, and the choice should be stated.
Small samples and bias in the effect size
With samples of 10 to 20, the standardized mean difference is biased upward unless corrected, which Hedges g does. Standard errors based on large-sample formulas are also unreliable, which affects the random-effects weights and confidence intervals. Methods that improve coverage when studies and samples are small, such as the Hartung-Knapp adjustment, help in the case of few studies. Methods to estimate heterogeneity can be unstable with few studies, and a review should say how many studies lie behind each moderator analysis. Because small studies are noisy, they are also most prone to selection for significance, which makes publication-bias analysis important in this field.
Multiple outcomes and time points
A training study may report strength on four lifts, tested at three times. Treating each as an independent study exaggerates precision. Multilevel models and robust variance estimation account for the dependence, as does selecting one outcome per study by a rule set in advance. Choosing the most favorable outcome after seeing results introduces bias. We prespecify the primary outcome category and time point, and we show sensitivity analyses using other choices. Where outcome measures differ (one-repetition maximum, isokinetic torque, jump height), they are analyzed by category and not combined into a single "strength" effect without justification.
Dose, intensity and moderators
Questions such as how many sets per week, what load or what interval length gives the best results are about dose-response. Direct comparisons of doses within trials are limited, so reviews often use meta-regression on study-level dose. This depends on the dose being defined consistently, which it often is not: intensity may be expressed as a percentage of maximum, as perceived exertion or as heart rate zones. Study-level associations are observational and affected by confounding among training variables, participants and duration. We use the dose-response methods described on the dose-response page for health questions and cautious meta-regression for training variables, and we describe results as associations.
Participants and generalizability
Training status, age, sex, health status and sport affect the response to exercise. Most studies recruit young, healthy men and university students, and women are under-represented in the literature on sport and exercise. Effects in trained athletes are generally smaller than in untrained participants because there is less room to improve, and a review should separate them. Results for older adults or patients are covered by other bodies of research, and we point to the relevant rehabilitation and clinical pages. Reviews should report the proportion of female participants and examine sex as a moderator when data allow.
Physical activity and health outcomes
Cohort studies relate self-reported physical activity to mortality, cardiovascular disease, diabetes and other outcomes, and dose-response meta-analyses have described falling risk with increasing activity, with the steepest gains at low levels. Self-reported activity is measured with error, and people who are healthier are more active, so reverse causation and confounding are concerns. Device-based measures, such as accelerometry, give better data and are increasingly used. A review should separate self-reported from device-measured activity, report adjusted estimates and state the covariates. The service does not give medical advice or exercise prescriptions.
Publication bias and selective reporting
Small studies with significant results are more likely to be published and to be reported with favorable outcomes selected. Few exercise studies are preregistered in a public registry, so comparing registered and reported outcomes is often impossible. Funnel plots, selection models and tests for small-study effects can show whether the evidence appears distorted. Since these tests are unreliable with few studies or high heterogeneity, we report several and avoid overinterpreting any one. Searching theses and conference abstracts, and writing to authors for unreported data, reduces the problem.
Supplements, recovery and nutrition studies
Many exercise trials test supplements (creatine, caffeine, protein, nitrate), recovery methods (cold water immersion, compression, sleep) or dietary strategies. These studies are often funded by manufacturers, use small crossover designs and report positive results on performance tests that are themselves variable from day to day. A review should record the funding source, the reliability of the performance test, and whether the participants and testers were blinded, which is possible for supplements with a matched placebo but rarely for recovery methods. Effects of recovery treatments on perceived soreness are measured subjectively and could reflect expectation. The review should separate subjective from objective outcomes and report the number of studies behind each.
Measurement reliability and the smallest worthwhile change
Performance tests vary from day to day because of biological and measurement variability. A change smaller than that variability cannot be told from noise in a single person, though averaging across a group still gives information. Sports scientists use the smallest worthwhile change, often defined as a fraction of the between-athlete standard deviation, to judge whether an effect would matter in competition. A review can report effects in natural units alongside standardized ones and relate them to published estimates of test reliability. The choice of threshold is a judgment and should be explained. Methods that classify effects as beneficial or trivial by comparing the confidence interval with a threshold have been criticized, and we report estimates with intervals and do not rely on such classifications.
These points matter for interpretation. A statistically significant gain of one centimeter in a jump test may be trivial in practice, while a nonsignificant gain from an imprecise study may include a worthwhile benefit. Showing both the estimate and the interval is the way to be fair to both readings, and it lets coaches and researchers apply their own thresholds to the same data.
Common pitfalls we look for
- Treating crossover trials as parallel-group or counting participants twice.
- Pooling pre-post changes with controlled comparisons.
- Using uncorrected standardized differences in very small samples.
- Counting multiple outcomes from one study as independent.
- Mixing trained and untrained participants.
- Interpreting study-level dose associations as causal.
Planning an exercise science synthesis
We help define participants, exercise, comparator and outcomes, specify the handling of crossover and pre-post designs, plan searches in MEDLINE, SPORTDiscus, Embase, CINAHL and Web of Science, and set up coding of training variables, participant characteristics and outcome measures. See the meta-analysis service for scope and process.
An invented example of a small-study problem
Suppose ten invented trials of a training method for sprint speed report an average standardized difference of 0.60, but the five smallest studies average 0.95 and the five largest 0.25. Funnel asymmetry like this suggests that small studies with large effects are overrepresented, perhaps through selective publication or through chance combined with selection. The honest summary is that the effect is probably nearer 0.25 than 0.60 and that the evidence is of low certainty. A review that reported 0.60 alone would encourage over-optimism.
Ethics and responsible interpretation
Exercise findings are widely quoted in the media and used by coaches and the public. A review states the populations studied and the limits of the evidence, and avoids language that suggests that a result applies to everyone who exercises. Where supplements are involved, conflicts of interest of the primary studies are recorded, since industry funding is common.
How we support research projects in this area
From small trials to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A question, effect sizes for repeated-measures designs, assumptions about correlations and a protocol for registration.
Searching and extraction
Searches of sport and health databases, with extraction from figures and checks between extractors.
Synthesis
Random-effects and multilevel models, dose-response and meta-regression, and small-study analyses.
Manuscript and submission
The manuscript, data and code, and journal preparation.
Boundaries of this service
An exercise science synthesis describes average effects across published studies. It does not provide training programs, coaching, medical advice or exercise prescriptions, and people with health conditions should seek advice from a qualified professional. Most studies are small and use young healthy participants, so results may not apply to other people.
Frequently asked questions
How are crossover trials handled?
With effect sizes that use the within-person correlation, assuming a value when unreported and testing sensitivity, and not treating them as parallel-group trials.
Why correct for small samples?
Uncorrected standardized differences are biased upward in small samples, so Hedges g and adjusted intervals are used.
Can pre-post studies be pooled with controlled trials?
Not in the same analysis. Without a control group, change may reflect factors other than the intervention.
How do you deal with many outcomes per study?
With multilevel models, robust variance estimation or a prespecified outcome, plus sensitivity analyses.
Does meta-analysis show the best training dose?
It can describe associations between dose and outcome, but study-level doses are confounded, so results are not prescriptions.
Do you provide exercise or training advice?
No. The service provides research and evidence-synthesis support only.
References
- Hedges LV. Distribution theory for Glass's estimator of effect size and related estimators. J Educ Stat. 1981;6(2):107-128.
- Morris SB, DeShon RP. Combining effect size estimates in meta-analysis with repeated measures and independent-groups designs. Psychol Methods. 2002;7(1):105-125.
- Pustejovsky JE, Tipton E. Meta-analysis with robust variance estimation: expanding the range of working models. Prev Sci. 2022;23(3):425-438.
- Schoenfeld BJ, Ogborn D, Krieger JW. Dose-response relationship between weekly resistance training volume and increases in muscle mass: a systematic review and meta-analysis. J Sports Sci. 2017;35(11):1073-1082.
- Arem H, Moore SC, Patel A, et al. Leisure time physical activity and mortality: a detailed pooled analysis of the dose-response relationship. JAMA Intern Med. 2015;175(6):959-967.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.