Meta-analysis in rehabilitation and physiotherapy

Rehabilitation research studies exercise, manual therapy, electrotherapy, education and multidisciplinary programs for people with injury, pain, stroke, heart and lung disease and other conditions. Trials are usually small, cannot blind participants and deliver variable doses of therapy. Reviews need to deal with these features and with outcome scales.

Evidence synthesis in rehabilitation

Rehabilitation aims to restore function and participation after injury, illness or surgery, or to help people live with long-term conditions. The Physiotherapy Evidence Database (PEDro) indexes thousands of trials and reviews and rates the quality of trials, and Cochrane Rehabilitation and related groups produce reviews on topics from exercise for low back pain to cardiac rehabilitation. The field now publishes many systematic reviews, and a recurring conclusion is that the evidence is of low certainty because trials are small and poorly reported.

Three features set this evidence apart. The intervention is behavioral and physical, with the therapist, the patient and the setting all affecting the result. Participants and therapists cannot be blinded, so subjective outcomes are vulnerable to bias. And outcomes are functional scales or performance tests with measurement error and debated minimal important differences. In addition, therapy is given in many forms, often as a package, with differences in dose that are inadequately reported.

Methods follow meta-analysis and systematic review, with the adaptations below. Related areas are orthopedics and neurology.

Describing and analyzing the dose of therapy

Exercise and other therapies have a dose: frequency, intensity, time and type (the FITT principles). Reports often omit details, and the programs delivered differ from those planned. The TIDieR checklist and the Consensus on Exercise Reporting Template (CERT) were developed to improve reporting. A review should extract these items, report their completeness and use them as moderators. Meta-regression of effect on total dose, weeks of therapy or supervision can show whether more therapy gives more benefit, though the relation is often unclear because dose is confounded with population and study design. Dose-response methods may describe non-linear relations, see dose-response meta-analysis.

Control conditions are varied: no treatment, usual care, sham, attention control or another active therapy. Comparisons against no treatment produce larger effects than comparisons against sham or other active treatments. Reviews should stratify by control type. A finding that exercise is better than no treatment says little about whether it is better than another therapy.

Outcome scales and the minimal important difference

Rehabilitation outcomes include pain scales, disability questionnaires, performance tests such as the six-minute walk test and the timed up-and-go, strength measures and quality-of-life instruments. When trials use different scales, the standardized mean difference is used, and it should be translated back to a familiar scale. Suppose a review finds a mean difference of 7 points on a 0 to 100 disability scale where the standard deviation is 18. The standardized difference is 0.39, a small effect by convention, and the difference is below a minimal important difference of 10 points. The numbers are invented. A statistically significant effect that is smaller than the minimal important difference is unlikely to be noticed by patients, and the review should say so. Reviews can also report the proportion of patients who achieve an important change if trial data allow.

Performance tests are measured by assessors who are often aware of allocation. Blinded assessment is possible and should be recorded. Instruments differ in responsiveness, and measurement properties can be checked with the COSMIN guidance.

Bias, small trials and reporting quality

Because blinding of participants and therapists is usually impossible, many rehabilitation trials are at high risk of performance bias for subjective outcomes. Placebo and sham techniques exist for some modalities, such as ultrasound, laser and taping, but not for exercise. Trials are small, often under 50 participants, and small-study effects appear in funnel plots. Reviews that find larger effects in smaller trials should note the possible reasons, including reporting bias, lower quality and differences in the intervention. The PEDro scale, a 10-item scale of trial quality, is commonly used in the field, though it does not replace a domain-based assessment such as RoB 2 because scores combine items of different importance.

Trial registration and reporting of outcomes are improving but remain inconsistent, and a comparison with registry entries often shows selective outcome reporting. Therapist allegiance can influence results: a trial run by developers of a technique may show larger effects than one run by independent investigators. Funding and investigator roles should be coded.

Complex and multidisciplinary programs

Cardiac, pulmonary and stroke rehabilitation, and chronic pain programs, combine exercise, education, psychological support and medical management. The effect of the package is the main question for services, but which element contributes is a question for science. Component analysis may help when trials vary elements independently; see component network meta-analysis. Qualitative studies on barriers and facilitators, such as access, motivation and the home environment, supply explanations, and a mixed-methods review can link them to effect sizes. Adherence and attendance are central: a program that people do not attend does not work, and reviews should extract participation rates. Telerehabilitation and home-based programs have grown; comparisons with center-based rehabilitation should report adherence, safety and equity of access.

Follow-up, maintenance and return to activity

Short-term benefit of rehabilitation often exceeds long-term benefit. Many trials end at the end of treatment, and those with follow-up show that effects fade when therapy stops. Reviews should present results at short, medium and long term, as described in longitudinal meta-analysis, and report the number of studies at each point. Return to work, sport or independent living are outcomes that matter to patients but are defined variably. Cost and resource use are also relevant to services, though economic evaluations alongside trials are often underpowered.

Condition-specific issues: musculoskeletal pain, stroke, cardiac and respiratory rehabilitation

Musculoskeletal pain conditions, such as low back and neck pain, have favorable natural histories, and many trials show similar improvement in treatment and control groups over time. The average difference between exercise and usual care at 3 months is small in many reviews, with larger effects for some forms of exercise and for supervised programs. Reviews must therefore separate short-term from long-term effects and record the type of exercise, since "exercise" covers strengthening, stabilization, aerobic and mind-body programs that differ in content. Education and graded activity are components of many programs and belong in the description.

Stroke rehabilitation trials deal with recovery that is greatest in the first months, so timing of therapy after onset matters, and outcomes such as walking speed and arm function change with the stage of recovery. Intensity of therapy and task specificity show consistent associations with outcome in meta-regressions. Cardiac and pulmonary rehabilitation have strong evidence for outcomes like exercise capacity, quality of life and hospital admissions in some groups, although participation in programs is low and uptake differs by sex, age and income. Reviews should record uptake and completion rates, since the benefit of a program is limited by how many people take part.

Across conditions, outcomes may be measured at the level of body function (strength, range of motion) or participation (return to work, social roles). Improvements in function do not always lead to improvements in participation. Reviews should keep these levels distinct, as the International Classification of Functioning suggests, and avoid concluding that a gain in a laboratory measure means a gain in daily life.

Interpreting results for clinicians and services

Clinicians need to know which therapy to offer, to whom, and with what expected benefit. A useful review reports effects on the original scale with the minimal important difference, the proportion likely to improve by an important amount, harms such as pain flare or falls, and the dose and setting used in the trials. Services need information on cost, adherence, staffing and whether the program can be delivered at scale or remotely. A review should state clearly what is unknown, for instance when trials include few older people or people with several conditions, which is common in rehabilitation, and when most trials come from a few countries with particular health systems. These statements allow readers to judge how far the evidence applies to their patients, and they point researchers to the trials that are most needed, such as pragmatic trials in community settings with the people who currently receive the least rehabilitation, including older adults and people in rural areas.

Common pitfalls we look for

  • Pooling different control types (no treatment, usual care, sham).
  • Describing therapy by its name only without dose and delivery.
  • Calling a significant effect meaningful without the minimal important difference.
  • Using the PEDro score as a total in place of domain judgments.
  • Ignoring attrition and adherence.
  • Presenting end-of-treatment results as long-term effects.

Planning and reporting

The protocol states the condition, the therapies and dose definitions, control types, outcomes with instruments and time points, and the analyses for dose, setting and population. Searches cover MEDLINE, Embase, CENTRAL, CINAHL, PEDro and registries. Risk of bias uses RoB 2, and the PEDro scale may be reported alongside it. Certainty is rated with GRADE, which often gives low certainty because of risk of bias and imprecision. Reporting follows PRISMA 2020, and the review protocol is registered in PROSPERO. Authors are encouraged to report dose with TIDieR and CERT items.

How we support research projects in this area

Support

From a clinical question to a published review

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A structured question, eligibility criteria and an analysis plan, with registration prepared where appropriate.

  • Searching and extraction

    Search strategies for the relevant databases and registries, screening and data extraction, and risk-of-bias assessment by design.

  • Synthesis

    Pairwise, network, diagnostic accuracy, prognostic or dose-response analysis, with a GRADE assessment for each outcome.

  • Manuscript and submission

    Reporting-guideline checklists, the manuscript and the preparation of submission materials.

Get a quoteDescribe your question, study types and target journal.

Boundaries of this service

A review of rehabilitation studies describes average effects in groups. It does not provide a treatment plan or advice on any person's rehabilitation, which depends on their condition, goals and the assessment of a qualified clinician. We do not prescribe exercise or interpret an individual's tests. People with pain, injury or illness should consult a physiotherapist or doctor.

Frequently asked questions

Why does the control group matter in rehabilitation reviews?

Effects against no treatment are larger than against sham or active therapy, so reviews should stratify by control type.

How do I describe an exercise intervention?

With frequency, intensity, time and type, and with TIDieR and CERT items, and use these as moderators.

What is the minimal important difference?

The smallest change in an outcome that patients perceive as important. Pooled effects should be compared with it.

Is the PEDro scale enough for risk of bias?

It is widely used, but a domain-based tool such as RoB 2 gives more informative judgments, and the two can be reported together.

How do I handle long-term follow-up?

Report results at short, medium and long term, with the number of studies at each point.

Do you provide treatment or exercise prescriptions?

No. The service provides research and evidence-synthesis support only.

References

  1. Hoffmann TC, Glasziou PP, Boutron I, et al. Better reporting of interventions: template for intervention description and replication (TIDieR) checklist and guide. BMJ. 2014;348:g1687.
  2. Slade SC, Dionne CE, Underwood M, Buchbinder R. Consensus on Exercise Reporting Template (CERT): explanation and elaboration statement. Br J Sports Med. 2016;50(23):1428-1437.
  3. Mokkink LB, de Vet HCW, Prinsen CAC, et al. COSMIN risk of bias checklist for systematic reviews of patient-reported outcome measures. Qual Life Res. 2018;27(5):1171-1179.
  4. Maher CG, Sherrington C, Herbert RD, Moseley AM, Elkins M. Reliability of the PEDro scale for rating quality of randomized controlled trials. Phys Ther. 2003;83(8):713-721.
  5. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.