Evidence synthesis in education
Meta-analysis is widely used in education. Syntheses of feedback, reading instruction, tutoring, formative assessment and technology use have shaped policy and practice. Some, such as Hattie's synthesis of syntheses, became widely known. They also attracted criticism for combining unlike studies, ranking effects without regard to the quality of the underlying evidence and using effect sizes as if they were comparable across designs and measures. These lessons apply to any review in the field.
Education studies have several features that matter for methods. Students are grouped in classes and schools, and randomization often happens at the class or school level. Outcomes are test scores, whose scale depends on the test, the age of the students and the time of year. Effects vary with the dose and quality of delivery, the teacher and the context. Many studies are small, short in duration and delivered by the developers of the program, which raises the risk of bias.
Our approach uses the framework in meta-analysis and systematic review, with the adjustments described below.
Effect sizes and how to read them
| Issue | Practice | Why it matters |
|---|---|---|
| Clustered assignment | Use effect sizes and standard errors adjusted for clustering (Hedges 2007) | Ignoring clustering makes intervals too narrow |
| Pre-test adjustment | Prefer adjusted post-test differences over raw gains where possible | Reduces baseline imbalance; mixing raw and adjusted values complicates pooling |
| Outcome measure type | Separate researcher-designed from standardized tests | Researcher-designed measures align closely with the intervention and give larger effects |
| Interpretation | Compare to benchmarks from the field, or to months of learning | A fixed 0.2 or 0.5 cut-off misrepresents typical effects in education |
| Sample size | Report effects by study size and design | Small studies show larger effects on average |
Kraft argued in 2020 that interpreting education effects against Cohen's benchmarks is misleading and proposed benchmarks based on the distribution of effects in rigorous studies, in which an effect of 0.05 to 0.20 standard deviations may be medium for causal studies with standardized outcomes. Bloom and colleagues similarly provided benchmarks based on typical annual growth in test scores. The conclusion for reviewers is to say how big an effect is in terms of a meaningful comparison, and to avoid applying one scale across different outcomes and ages.
Clustered designs and multiple outcomes
A trial that randomizes schools should be analyzed with the school as the unit, or with an adjustment that uses the intraclass correlation. Reviews depend on what authors report. If a study analyzed students as if they were independent, its standard error is too small, and the review should adjust it using a stated or assumed intraclass correlation. Hedges gave formulas for this in 2007. The assumption about the intraclass correlation should be varied in a sensitivity analysis, since the correction can be large.
Studies also report several outcomes, such as reading comprehension and vocabulary, and several time points, such as immediate and delayed post-tests. As in other fields, these dependent effect sizes require a multilevel model or robust variance estimation. Hedges, Tipton and Johnson described the latter, and it is widely used in education reviews because the correlation between outcomes is rarely known.
Moderators: dose, delivery and context
Education interventions vary in many ways, and the main use of meta-regression is to relate effect size to features of the intervention and the study. Typical variables include the age of students, subject area, length and intensity of the program, who delivered it, group size, whether the developers conducted the evaluation and the type of outcome measure. Studies by developers tend to show larger effects than independent evaluations, and the review should test for this.
Moderator analysis has to be treated with caution. Features of interventions are chosen by researchers and are not randomly assigned across studies, so they are confounded with each other and with study quality. A finding that longer programs show larger effects might be due to differences in the outcome measure, not to duration. Presenting such findings as hypotheses unless they come from planned analyses, or from studies that compare features within a trial, is an honest approach.
Best-evidence synthesis and alternatives to pooling
Slavin proposed best-evidence synthesis in 1995 as an alternative to indiscriminate pooling. It applies explicit standards for the quality of studies (for instance, a minimum duration, adequate control groups and independent outcome measures), includes only those that meet them and presents effect sizes in the context of the details of each study. The rationale is that including studies of poor quality or short duration inflates effects. A comparison of results with and without the lower-quality studies shows how much depends on the inclusion rule.
When programs and outcomes are too varied to pool, a systematic review with a narrative synthesis, a scoping review or an evidence map may fit better. Realist reviews, described under realist review, aim to explain how and why programs work in different contexts. The choice should follow the question the review wants to answer.
Bias, grey literature and reporting
Education has a large grey literature of reports, dissertations and evaluations, and a review that restricts itself to journals risks publication bias. Searches include ERIC, PsycINFO, Education Source, Scopus and Web of Science, plus repositories such as the What Works Clearinghouse and the Education Endowment Foundation, and theses databases. Funnel-plot-based tests should be interpreted with caution because small studies may differ in design as well as in selective reporting. Reporting follows PRISMA 2020, and Campbell Collaboration guidance for reviews in education and related fields provides further detail. Protocols can be registered with the Campbell Collaboration or the Open Science Framework, and the coding sheet, the list of excluded studies with reasons and the analysis code should be kept so that the review can be updated when new trials appear. Education evidence ages quickly as curricula, technology and tests change, so the date of the last search should be stated clearly, and updates planned if the topic is active. Readers can then judge whether the review still reflects current practice and current tests.
Specialties and sub-fields
Sub-fields differ in populations, outcomes and designs. Pages for sub-fields are added as they are completed.
Higher education
Teaching methods, retention and student outcomes in colleges and universities.
STEM education
Science, technology, engineering and mathematics instruction.
Medical education
Simulation, assessment and training of health professionals.
E-learning
Online, blended and technology-enhanced learning.
Special education
Interventions for learners with disabilities, with single-case and group designs.
A worked adjustment for clustering
Consider a trial that assigned whole classrooms to a reading program, with 600 students in classes of about 25. The outcome was analyzed as if the 600 students were independent. Suppose the intraclass correlation is 0.20, a value that is plausible for reading scores and is used here only to illustrate the arithmetic. The design effect is 1 + (m - 1) x rho = 1 + 24 x 0.20 = 5.8. The effective sample size is 600 / 5.8, which is about 103 students.
The trial has the information of roughly 103 independent students, not 600. If its unadjusted standard error were used in a meta-analysis, the study would carry about 5.8 times too much weight, and the pooled interval would be too narrow. The review should therefore divide the study's weight by the design effect, or use an effect size and variance computed with the cluster adjustment. Repeating the calculation with an intraclass correlation of 0.10 and 0.30 shows how sensitive the result is; this is the sensitivity analysis that a reader should expect.
Interpreting results for practice
Teachers and school leaders want to know whether a program is worth the cost and effort. A pooled effect size is a start. A report is more useful when it states the effect in terms of months of learning or percentile movement, the range of effects across settings, the conditions that seemed to matter, the cost and the quality of the evidence. Interpretation should take account of the age of the students, since a standard deviation represents more learning in younger children than in older ones.
Fidelity is important. An intervention works as designed only when it is delivered as designed. Reviews can code whether studies reported fidelity checks, training and support for teachers, and compare effects. Programs that need heavy support may show good effects in studies that supply it and weaker effects when scaled up. Studies of scaled-up programs show smaller effects on average than small efficacy trials, and a review that separates the two gives a more realistic picture.
Single-case and small-sample designs
Special education and some behavioral interventions rely on single-case experimental designs, in which a few participants are observed repeatedly under baseline and intervention conditions. Meta-analysis of such designs is possible, with effect sizes such as the non-overlap of all pairs, the percent exceeding the median or, more recently, standardized mean differences designed for single-case data and multilevel models. Each has limits, and results from different metrics may not agree. Reviews of single-case research should therefore state the design standards used to screen studies, describe the metric and its weaknesses, and combine visual analysis with the statistics. They should not be mixed with group-design effect sizes without a clear justification, since the two measure different things.
How we support research projects in this area
From a research question to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A question, inclusion standards for study quality, and a plan for clustered designs and outcome types.
Searching and extraction
Searches of education databases and repositories of evaluations, with double coding of effect sizes and moderators.
Synthesis
Random-effects and robust variance models, adjustment for clustering, moderator analysis and sensitivity analysis.
Manuscript and submission
The manuscript, coded-study tables and journal preparation.
Boundaries of this service
A meta-analysis summarizes studies of groups of learners. A pooled effect does not tell a teacher what will work with a particular student or class, and effects depend on how a program is delivered. We do not provide advice on the education of an individual learner. The result of a review informs decisions; it does not make them.
When studies are too different to combine, we recommend a systematic review or evidence map and say so at the start.
Frequently asked questions
How should I handle studies that randomize classrooms or schools?
Use effect sizes adjusted for clustering, with a stated intraclass correlation if the study did not report one, and test the assumption in a sensitivity analysis.
Is an effect size of 0.2 small in education?
Not necessarily. In rigorous education trials with standardized outcomes, effects of 0.05 to 0.20 are common, so field-specific benchmarks are more informative than Cohen's cut-offs.
Should I combine researcher-designed and standardized tests?
Usually not without comment. Researcher-designed measures tend to show larger effects, so the two are analyzed separately or tested as a moderator.
What is best-evidence synthesis?
A method that includes only studies meeting stated quality standards and presents effect sizes with study details, instead of pooling all available studies.
Can you include grey literature?
Yes. Searches can include repositories of evaluations, reports and dissertations, subject to the scope agreed.
Do you advise on individual students?
No. The service covers research and evidence-synthesis support only.
References
- Hedges LV. Effect sizes in cluster-randomized designs. J Educ Behav Stat. 2007;32(4):341-370.
- Kraft MA. Interpreting effect sizes of education interventions. Educ Res. 2020;49(4):241-253.
- Bloom HS, Hill CJ, Black AR, Lipsey MW. Performance trajectories and performance gaps as achievement effect-size benchmarks for educational interventions. J Res Educ Eff. 2008;1(4):289-328.
- Slavin RE. Best evidence synthesis: an intelligent alternative to meta-analysis. J Clin Epidemiol. 1995;48(1):9-18.
- Hedges LV, Tipton E, Johnson MC. Robust variance estimation in meta-regression with dependent effect size estimates. Res Synth Methods. 2010;1(1):39-65.
- Hattie J. Visible learning: a synthesis of over 800 meta-analyses relating to achievement. London: Routledge; 2009.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.