What comparative effectiveness research is
Many trials compare a treatment with placebo or with nothing. That design shows whether an option works, but it does not tell a clinician or a health system which of several active options to choose. Comparative effectiveness research addresses that gap. The Institute of Medicine in the United States described it in 2009 as the generation and synthesis of evidence that compares the benefits and harms of alternative methods to prevent, diagnose, treat and monitor a condition, with the purpose of helping people make informed decisions.
The definition has two parts that matter for reviews. The first is the comparison of active alternatives, not just against placebo. The second is the focus on decisions in practice, which pushes attention toward diverse populations, usual-care settings and outcomes that matter to patients. Evidence synthesis supports both parts. It combines the available comparisons, and it asks how well the evidence applies to the people the decision concerns.
The field is not limited to medicine. Education, social policy and health services research ask the same question of programs: which teaching method, which housing support, which care model. The same principles apply, although the designs and outcomes differ.
Evidence sources and what each can show
A comparative effectiveness question rarely has one ideal study. The usual situation is a mixed body of evidence, and a review needs to say what each type contributes.
| Source | What it provides | Main strength and limit |
|---|---|---|
| Randomized trials | Head-to-head or pragmatic trials of alternatives | Strongest protection against confounding; may be selective and short |
| Network meta-analysis | Indirect and mixed comparison across a connected set of trials | Compares treatments never tested directly; depends on transitivity |
| Cohort studies | Observational comparison in routine care | Broad populations and long follow-up; confounding is the main threat |
| Registries and routine data | Large databases of practice | Scale and relevance; data quality and missing information vary |
| Target trial emulation | Design the observational analysis to mimic a trial | Clarifies eligibility, time zero and treatment strategies; does not remove unmeasured confounding |
Head-to-head randomized trials are the preferred evidence but are often missing. Pharmaceutical trials usually compare a new product with placebo or with an older standard, so the network of evidence is built around common comparators. Pragmatic trials, which enroll broader populations and use usual care as the comparison, answer questions about effectiveness in routine practice, but they are expensive and less common.
Indirect comparison and network meta-analysis
When two treatments have both been compared with a third, their relative effect can be estimated indirectly. Network meta-analysis extends this to many treatments, combining direct and indirect evidence into a consistent set of estimates and, if wanted, a ranking. This is the main statistical route to comparative effectiveness when head-to-head trials are scarce.
The method rests on transitivity: the trials comparing different pairs must be similar enough in the factors that modify the effect that an indirect comparison is fair. If trials of treatment A against placebo enrolled mild patients and trials of treatment B against placebo enrolled severe patients, the indirect A versus B estimate is biased. Reviewers check transitivity by tabulating effect modifiers across comparisons and testing for inconsistency between direct and indirect evidence, while remembering that those tests have low power.
Rankings should be handled with care. A treatment can be ranked first with a very wide interval. The certainty of each comparison should be rated, and approaches such as CINeMA are designed for that. A ranking without the certainty behind it can mislead a reader more than no ranking at all.
Observational evidence and confounding
Randomized trials often exclude older people, people with several conditions and people on other medicines, so the results may not apply to the patients seen in practice. Observational studies include these people, which makes them attractive for comparative effectiveness, and they can follow patients longer. The cost is confounding. People who receive one option usually differ from people who receive another in ways that also affect the outcome, such as severity, age or access to care.
Reviews of observational comparisons should extract adjusted estimates and note the covariates in each model, assess the risk of bias with a tool such as ROBINS-I, and examine whether the direction and size of results differ between designs. Pooled observational estimates inherit the confounding of their components and do not cancel it; meta-analysis makes the estimate more precise but not more valid. A narrow interval around a biased estimate is a risk, not a comfort.
Target trial emulation has become a common framework for the primary studies. The analyst specifies the trial that would answer the question, including eligibility, treatment strategies, assignment, follow-up start, outcomes and analysis, and then designs the observational analysis to match. This helps avoid errors such as immortal time bias. It still depends on measured confounders, and reviews should check how well each study followed the framework.
Applicability, heterogeneity and who benefits
An average effect may not describe any individual. Comparative effectiveness aims to say which option suits which group, so subgroup and moderator analyses matter. They are also easy to misuse. The review should pre-specify a small number of subgroup questions on biological or contextual grounds, test the interaction and not the within-subgroup significance, and present findings from subgroup analysis as hypotheses unless the evidence is strong.
Applicability is a separate judgment from internal validity. GRADE calls it indirectness: whether the populations, interventions, comparators and outcomes in the studies match those in the question. A comparison supported by studies of younger patients in specialist centers is indirect evidence for older patients in community clinics. Stating the differences openly, and lowering the certainty when they matter, helps the reader apply the result.
Absolute effects deserve as much attention as relative ones. A relative risk of 0.80 means different things when the baseline risk is 2 percent and when it is 30 percent. Converting relative effects to absolute differences for stated baseline risks is part of good reporting.
Outcomes that matter to decisions
The choice of outcome shapes the conclusion. A comparison of two treatments for blood pressure on a laboratory marker may favor one, while their effect on cardiovascular events is unclear. Reviews should prioritize outcomes that patients and clinicians consider important, include harms as well as benefits and be explicit about surrogate outcomes and the strength of their link to clinical outcomes.
Core outcome sets, where they exist, help to align the outcomes across studies and reduce the problem of heterogeneous measurement. When a review compares options with different benefit and harm profiles, a table that sets benefits and harms side by side, with certainty ratings, serves a decision maker better than a single summary score.
Common pitfalls
Several errors recur in comparative effectiveness reviews. The first is treating the network as a bag of trials, without describing which comparisons exist. A network graph that shows the number of trials and participants on each edge lets the reader see where the evidence is thin and which estimates rest on a single small trial.
The second is mixing populations or doses without comment. If one treatment was tested at low dose in mild disease and another at high dose in severe disease, the numerical comparison reflects the populations and the doses as much as the treatments. Splitting nodes by dose, or restricting the network to a defined population, is often more honest than lumping.
The third is industry influence. Trials sponsored by a manufacturer tend to favor the sponsor's product, and the review should record funding, include it in risk-of-bias judgment and examine whether results differ by sponsor. The fourth is an unstated change of question after results are known. Registered protocols, with dated amendments, guard against this. Finally, avoid presenting a statistically significant difference as a clinically important one; compare the estimate with a stated threshold of importance.
Planning a comparative effectiveness review
A review for this purpose begins with a protocol that states the decision problem. It defines the population, the set of alternatives (which options are in the network and why), the outcomes and the types of study that will be included. It declares in advance how randomized and observational evidence will be handled: pooled together, analyzed separately or used in a hierarchy. It registers the review and plans the assessment of transitivity before the data are seen.
Stakeholder input is an accepted part of the field. Patients, clinicians and policy makers can help to define options and outcomes that matter. The review team should record how this input shaped the question and avoid changing the question once results are known.
Reporting follows PRISMA 2020 and its extension for network meta-analysis, with the GRADE approach for certainty. Results are best presented with a network graph, relative and absolute effect tables, a statement of the assumptions that cannot be tested and, where used, rankings with their uncertainty.
Limitations
A comparative effectiveness review is limited by the evidence available. Head-to-head trials are often missing, indirect comparisons depend on assumptions that cannot be fully tested and observational studies carry confounding that pooling does not remove. Rankings can overstate differences. Results from trial populations may not apply to the people a decision concerns, and the evidence for harms and long-term outcomes is often weaker than for short-term benefit. The review reports what the evidence supports and rates its certainty; it does not choose a treatment for a person.
Evidence synthesis informs decisions about groups and policy. It does not replace professional judgment about an individual.
How we can help
Support for comparative effectiveness research
The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.
Feasibility check
A review of your studies and data to confirm that the method is suitable and which approach fits.
Analysis and figures
The analysis run to a prespecified plan, with forest plots and the other figures.
Methods and results text
Written for the manuscript and aligned with PRISMA 2020 or the relevant extension.
Manuscript and submission
Optional: the full paper, the reporting checklist and the submission materials.
Frequently asked questions
What is comparative effectiveness research?
Research that compares the benefits and harms of alternative options in practice to help people choose between them.
How is it different from a standard efficacy trial?
Efficacy trials often test an option against placebo under controlled conditions. Comparative effectiveness studies compare active alternatives in settings that resemble routine care.
Can meta-analysis compare treatments that were never compared directly?
Yes. Indirect comparison and network meta-analysis estimate such comparisons, provided the trials are similar enough in effect modifiers (transitivity).
Can observational studies be combined with trials?
They can, but they answer the question with different biases. Many reviews analyze them separately and compare the results instead of mixing them.
Do rankings tell me which treatment is best?
Not by themselves. A ranking can hide wide uncertainty, so it should appear with the effect estimates and the certainty of the evidence.
Why report absolute effects?
Because a relative effect has a different practical meaning depending on the baseline risk of the group.
References
- Sox HC, Greenfield S. Comparative effectiveness research: a report from the Institute of Medicine. Ann Intern Med. 2009;151(3):203-205.
- Salanti G. Indirect and mixed-treatment comparison, network, or multiple-treatments meta-analysis: many names, many benefits, many concerns for the next generation evidence synthesis tool. Res Synth Methods. 2012;3(2):80-97.
- Hutton B, Salanti G, Caldwell DM, et al. The PRISMA extension statement for reporting of systematic reviews incorporating network meta-analyses of health care interventions. Ann Intern Med. 2015;162(11):777-784.
- Sterne JA, Hernan MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
- Hernan MA, Robins JM. Using big data to emulate a target trial when a randomized trial is not available. Am J Epidemiol. 2016;183(8):758-764.
- Tunis SR, Stryer DB, Clancy CM. Practical clinical trials: increasing the value of clinical research for decision making in clinical and health policy. JAMA. 2003;290(12):1624-1632.
- Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082.