Evidence synthesis in medical education
Medical education research covers selection of students, curriculum, teaching methods, simulation, assessment, feedback, professionalism, continuing education and the learning environment. Systematic reviews are common in this field, and the Best Evidence Medical Education collaboration has produced many. Meta-analyses have been published on simulation-based training, online learning for health professionals, feedback, and assessment tools.
The field has particular features. Experiments are small and conducted in single institutions. Control conditions are often "no intervention" or poorly specified. Outcomes are most often satisfaction, knowledge or skill measured in a simulated setting, and only rarely behavior in practice or patient outcomes. Educational interventions are complex, and context matters a great deal. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for education and connects with the medical and health sciences area.
Outcome levels
| Level | Example | Issue for synthesis |
|---|---|---|
| Reaction | Satisfaction, perceived usefulness | Easy to measure, weakly related to learning |
| Learning | Knowledge tests, skill ratings in simulation | Most common outcomes; measures are often study-specific |
| Behavior | Performance with real patients, observed in practice | Rare, expensive; rater bias and variability |
| Results | Patient outcomes, system outcomes | Very rare; confounded by many other factors |
Reviews should report outcome level and avoid summarizing results across levels. A favorable effect on satisfaction does not show an effect on practice. Where studies report outcomes at several levels, we extract each. We also note how many studies report each level, since a conclusion about "effectiveness" based mostly on satisfaction and simulated skill says little about care.
Simulation-based education
Meta-analyses of technology-enhanced simulation have found large average effects on knowledge and skills in simulated settings compared with no intervention, smaller effects compared with another active instruction, and fewer data on patient outcomes. A large effect against no training is not surprising, and it says little about whether the simulation adds value over other methods. A review should separate comparisons with no intervention from comparisons with alternative instruction, and describe what differed between the active conditions. Features such as repetitive practice, feedback, range of difficulty and distribution of practice over time have been associated with larger effects, and these can be coded and examined.
The fidelity of the simulator (how closely it resembles reality) matters less than often assumed in many studies, and a review should look at the evidence rather than assume it.
Assessment tools and validity evidence
Reviews of assessment look at reliability and validity of tools such as objective structured clinical examinations, workplace-based assessments and multiple-choice tests. The standard framework for validity looks at content, response process, internal structure, relationships with other variables and consequences. Reliability coefficients can be pooled after transformation, and heterogeneity is large since reliability depends on sample, number of items or raters, and context. A review should avoid giving a tool a single reliability value without showing the conditions. Predictive validity studies, for example of selection tests and later performance, usually find modest correlations, often reduced by restriction of range, which a review can address with corrections while reporting the uncorrected values.
Heterogeneity, context and what works for whom
Educational interventions work through mechanisms that depend on learners, teachers and settings. Pooled effect sizes with very high heterogeneity are common, and the pooled value tells little about any one setting. Realist review and qualitative synthesis help explain how and why interventions work in some contexts, and a mixed approach can combine pooled estimates with explanations. We consider realist methods when the question concerns mechanisms, and describe them on the page for realist review.
Some meta-analyses in this field have been criticized for pooling very different interventions and comparators. We guard against this by defining the intervention carefully, grouping by comparator and showing the distribution of study-level effects in addition to the average.
Study design and reporting quality
Many studies are single-group pre-post or non-randomized. Reviews of the field have found that reporting is often poor, with missing details about the intervention, learners, outcome measures and randomization. Tools such as the Medical Education Research Study Quality Instrument assess aspects of design, sampling, outcomes and analysis. We use such tools, report them by study and test whether study quality relates to effect size. Reporting guidelines for educational interventions, including GREET, help describe interventions in a way that allows replication, and we use them to code what the intervention contained.
Workplace learning, feedback and continuing education
Studies of feedback, coaching, audit and continuing professional education look at changes in clinicians' practice. Reviews of audit and feedback in healthcare have found small to moderate average improvements in professional practice, with variation explained in part by the baseline performance, the source and format of feedback, and whether it came with an action plan. These findings are summarized in Cochrane reviews and can be linked with the clinical reviews in the medical and health sciences area. Education reviews should take care to separate changes in clinician behavior from changes in patient outcomes, which are rarely measured.
Publication bias and enthusiasm bias
Educators who design a new method usually evaluate it themselves in their own institution, and positive results are more likely to be published than null ones. We check small-study effects, search conference abstracts and compare developer-led evaluations with independent ones. Medical education has a tradition of reporting satisfaction results, which are almost always positive; they should not be taken as evidence of effectiveness on their own, and reviews should report them in a separate column.
Selection, admissions and predictive validity
Studies of medical school and residency selection relate admission measures, such as prior grades, aptitude tests and interviews, to later performance on examinations and, less often, in practice. Correlations are modest and are reduced when the studied group has already been selected on the predictor, which is called restriction of range. Corrections for range restriction and unreliability raise the estimates but depend on assumptions about the applicant pool that are often not reported. A review should give uncorrected and corrected values and state the source of the assumptions. It should also report the criterion used. Predicting later examination scores is easier than predicting how someone will practise, and results for the first should not be described as if they applied to the second.
Fairness is an important question in selection. Reviews can examine whether predictors function differently across groups, but evidence is limited, and a synthesis should be careful to report what the studies measured and to avoid conclusions about individuals. Evidence on interviews and situational judgment tests is mixed, and the reliability of unstructured interviews is typically lower than that of structured formats, which a review can report with the number of studies behind each statement.
Coding and transparency
Coding frames for medical education reviews record the profession and stage of training, specialty, country, class size, intervention features (duration, intensity, feedback, instructor), comparator, outcome level, instrument and its validity evidence, design, and the funding source. Two coders work independently on a sample, agreement is reported and the coded data set and analysis code are shared. Because intervention descriptions are often brief, we contact authors when details that affect coding are missing, and we record when no answer was received. We also record whether a study was part of a larger program of work by the same group, to avoid counting related reports as independent.
Common pitfalls we look for
- Pooling comparisons against no intervention with comparisons against active alternatives.
- Summarizing satisfaction, knowledge and patient outcomes together.
- Treating simulator fidelity as an explanation without evidence.
- Quoting a single reliability value for an assessment tool.
- Ignoring restriction of range in predictive validity studies.
- Assuming results from one school or country apply elsewhere.
Planning a medical education synthesis
We help define the learners, the intervention, the comparator and the outcome level, plan searches in MEDLINE, ERIC, Embase, CINAHL and PsycINFO, and set up coding of intervention features, design and context. Where the question concerns validity of assessments, we plan extraction of reliability and validity evidence in line with a recognized framework. See the systematic review service for scope and process.
An invented example of interpreting simulation results
Suppose a review finds a pooled standardized difference of 0.9 for simulation training against no intervention on simulated skill, 0.2 against traditional training, and no study measuring patient outcomes. The first number is large, but it answers a weak question. The second suggests that simulation adds a modest amount over existing teaching, with the added cost to be weighed. The absence of patient outcome data means that the review cannot say that patients benefit. A summary that reported only the first number would be misleading, so the review should present all three facts together.
Ethics and trainee data
Education studies involve trainees whose performance data are sensitive and sometimes linked to career consequences. A review uses published aggregate data and does not need identifiable trainee data. We note where primary studies reported ethical approval and consent and where they did not, and we do not include any report that identifies individual trainees or programs in a way that could harm them.
How we support research projects in this area
From a research question to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A question, inclusion standards for study quality, and a plan for clustered designs and outcome types.
Searching and extraction
Searches of education databases and repositories of evaluations, with double coding of effect sizes and moderators.
Synthesis
Random-effects and robust variance models, adjustment for clustering, moderator analysis and sensitivity analysis.
Manuscript and submission
The manuscript, coded-study tables and journal preparation.
Boundaries of this service
A medical education synthesis describes published evidence on training and assessment. It does not accredit programs, assess any trainee or provide clinical training advice. Many studies measure satisfaction or simulated skills in single institutions, so claims about effects on patient care require evidence that is often absent.
Frequently asked questions
Is simulation-based training effective?
Against no training, reviews find large effects on knowledge and skills in simulated settings. Against other training, effects are smaller, and patient outcomes are rarely measured.
Why separate outcome levels?
Satisfaction, learning, behavior in practice and patient outcomes answer different questions and should not be combined.
Can assessment reliability be pooled?
After transformation, with attention to how it depends on items, raters and sample. A single value for a tool is misleading.
How do you handle very high heterogeneity?
By defining interventions and comparators carefully, grouping them, showing study-level effects and, where useful, using realist or qualitative synthesis for explanations.
Are satisfaction results useful?
As descriptions of learner reaction, yes. As evidence of effectiveness, no.
Do you accredit programs or assess trainees?
No. The service provides research and evidence-synthesis support only.
References
- Cook DA, Hatala R, Brydges R, et al. Technology-enhanced simulation for health professions education: a systematic review and meta-analysis. JAMA. 2011;306(9):978-988.
- Cook DA, Brydges R, Zendejas B, Hamstra SJ, Hatala R. Technology-enhanced simulation to assess health professionals: a systematic review of validity evidence, research methods, and reporting quality. Acad Med. 2013;88(6):872-883.
- Reed DA, Cook DA, Beckman TJ, Levine RB, Kern DE, Wright SM. Association between funding and quality of published medical education research. JAMA. 2007;298(9):1002-1009.
- Ivers N, Jamtvedt G, Flottorp S, et al. Audit and feedback: effects on professional practice and healthcare outcomes. Cochrane Database Syst Rev. 2012;(6):CD000259.
- Phillips AC, Lewis LK, McEvoy MP, et al. Development and validation of the guideline for reporting evidence-based practice educational interventions and teaching (GREET). BMC Med Educ. 2016;16:237.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.