Evidence synthesis in special education
Special education research asks how to teach learners with disabilities, learning difficulties and developmental differences, what supports help them succeed, and how inclusion in ordinary classrooms works. Topics include reading and mathematics instruction for students with learning disabilities, behavioral and social-skills interventions, communication supports, transition to adulthood and the effects of placement. Meta-analyses have summarized explicit instruction, strategy instruction, peer-mediated approaches, and interventions for students on the autism spectrum.
The evidence has features that call for care. Participants are diverse and numbers per study are small. Many studies use single-case experimental designs, which are strong for within-person causal inference but yield data that standard meta-analysis cannot handle directly. Labels such as learning disability are defined differently across countries and periods, which limits comparability. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for education.
Single-case designs and their effect sizes
| Measure | What it does | Issue for synthesis |
|---|---|---|
| Percentage of non-overlapping data (PND) | Share of intervention points above the highest baseline point | Simple but sensitive to one extreme baseline point; no sampling variance; ceiling effects |
| Tau-U | Non-overlap with correction for baseline trend | Widely used; variance properties debated; complex calculation |
| Log response ratio | Log of ratio of means for a behavior | Suits outcome measured in counts or percentages; has a sampling variance |
| Between-case standardized mean difference | Effect across cases comparable to group-design SMD | Needs multiple-baseline designs with enough cases; assumptions about the model |
| Visual analysis | Judging level, trend, variability and immediacy of change | Standard in the field; subjective but transparent when documented |
Single-case studies repeatedly measure each participant before and during an intervention and show effect by replication within and across participants. They are central in special education because randomizing large groups is often impractical. Synthesizing them calls for effect sizes that suit the data, and for multilevel models that treat observations as nested in cases and cases in studies. A review should state the metric and software, show visual summaries of the studies, and report quality against recognized standards such as those of the What Works Clearinghouse for single-case designs. Results from single-case and group designs are best reported separately, with a comparison where the questions match.
Describing participants
Participants in special education studies differ in the type and severity of disability, age, language, additional conditions and prior supports, and studies describe them using different criteria. Some use a school-assigned label, others use test cutoffs, and others report only that students were "at risk". A review should record how each study identified participants and avoid combining different groups without a moderator. Subgroup comparisons based on small numbers should be treated as exploratory. Studies seldom report race, language and socioeconomic status in enough detail to examine equity, which is itself a finding.
Language matters. We use terms that are accurate and respectful, in line with current usage in the field, and we state the labels used by the original studies when they differ. We avoid language that implies deficits as the cause of educational outcomes where the cause lies in the environment.
Instruction and intervention effects
Reviews of reading and mathematics interventions for students with learning disabilities have found that explicit, systematic instruction, strategy instruction and the use of visual representations tend to produce positive effects, with larger effects on measures close to the taught content. Studies commonly use researcher-designed measures, small groups and instruction delivered by the research team, so effects in ordinary classrooms may be lower. A review should code the interventionist, the group size, the dosage and the type of measure, and compare effects on near-transfer with those on distant standardized tests. Follow-up measures after instruction has ended are less often reported and reveal whether effects last.
Inclusion and placement
Research on inclusion compares outcomes for students in general classrooms with those in separate settings, and studies of supports that make inclusion work. Comparisons are heavily confounded, because students placed in separate settings usually have greater needs, and statistical adjustment cannot fully correct it. Reviews have therefore been cautious in drawing causal conclusions about placement. Studies of the experiences of students, families and teachers add information about belonging and barriers. We combine such qualitative studies with quantitative results through mixed-methods or qualitative synthesis, and describe the limits plainly.
Autism and behavior interventions
Reviews of interventions for autistic learners have examined early behavioral approaches, social skills groups, communication supports, peer mediation and technology-based aids. Evidence quality varies, many studies are small or single-case, and reviews have classified practices by the strength of evidence under frameworks developed for the field. Outcomes reported by parents or therapists who know the intervention are open to bias, and those assessed by blinded raters are more credible. Neurodiversity-informed perspectives question which outcomes are valued; reviews should say who selected outcomes and whether learners' own views were taken into account. The service does not give clinical advice or recommend an intervention for any child.
Publication bias in single-case and group studies
Single-case studies are often published when results are clear, and publication bias in this literature has been shown by comparing published studies with dissertations. Funnel plots and selection models for single-case effect sizes are still developing, so we describe the evidence and use available methods with caution. We search dissertations and conference papers, and note differences between sources.
Transition, communication and behavior supports
Beyond academic skills, special education research addresses transition from school to adult life (employment, further study, independent living), communication supports including augmentative and alternative systems, and positive behavior support at individual and school level. Transition studies often use follow-up surveys with high loss of participants, and outcomes such as employment depend on local labor markets. Communication studies are mostly single-case with very few participants. School-wide behavior support has been evaluated in cluster trials with measurable effects on office referrals and school climate, and these are the main source of group-design evidence in this area.
For each of these topics, a review reports the number of studies, the number of participants and the designs, so that readers can see how thin the evidence can be. When fewer than a handful of studies address a question, a narrative description is more honest than a pooled estimate. We also look for studies that involved learners, families and teachers in setting outcomes or interpreting results, because the choice of which outcomes matter is not only a technical matter.
Fidelity, dosage and implementation
Interventions in special education are often delivered by teachers or aides after short training, and the dose actually received can be well below what was planned. Studies that report fidelity measures allow a review to ask whether higher fidelity goes with larger effects, a question that has an answer only if fidelity is measured and reported consistently. We code whether fidelity was measured, how, and what level was reached, and we examine dosage in terms of sessions and total minutes. Where implementation data are missing, we say so, since an apparently ineffective intervention may be one that was never delivered as intended.
Common pitfalls we look for
- Using a non-overlap index alone as if it were a standardized effect.
- Pooling single-case and group studies without separation.
- Combining participants with different disabilities without a moderator.
- Treating placement comparisons as causal.
- Relying on outcomes rated by people who delivered the intervention.
- Using deficit-based language that the studies themselves do not support.
Planning a special education synthesis
We help define the population, the intervention and the outcomes, decide how single-case and group studies will be handled, plan searches in ERIC, PsycINFO, Scopus and dissertation databases, and set up coding of participant description, design, interventionist, dosage and outcome type. See the systematic review service for scope and process.
An invented example
Suppose a review of a reading strategy intervention finds an average standardized difference of 0.55 in group studies on researcher-designed tests, and 0.15 on standardized tests, and that in 20 invented single-case studies most participants show clear improvement. These results fit together: the intervention seems to help students with the taught skills and to transfer little to general reading scores. The review would report both kinds of study, state the number of participants in each, and say that the evidence does not show how well the approach works in a class with typical staffing. It would not state that the approach "works" or "does not work" in general.
Ethical considerations
Studies involve children and young people who are vulnerable, and consent and assent procedures are important. Reviews use published data and should note where reports gave identifiable details of single participants. In single-case studies, a few participants may be identifiable from descriptions, so extracted data are limited to what is needed for synthesis, and shared data sets do not include identifying information.
Coding and transparency
Coding frames record disability category and how it was defined, age, grade, setting, interventionist, dosage, group size, design, quality indicators, outcome type and whether a measure was taught or generalized. For single-case studies we extract data points from graphs using software and report the reliability of extraction by having two coders extract a sample independently. We share the coded data and analysis code with the final report.
How we support research projects in this area
From a research question to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A question, inclusion standards for study quality, and a plan for clustered designs and outcome types.
Searching and extraction
Searches of education databases and repositories of evaluations, with double coding of effect sizes and moderators.
Synthesis
Random-effects and robust variance models, adjustment for clustering, moderator analysis and sensitivity analysis.
Manuscript and submission
The manuscript, coded-study tables and journal preparation.
Boundaries of this service
A special education synthesis describes average effects across published studies. It does not diagnose, determine eligibility, write individual education plans or recommend any intervention for a particular child. Samples are small and diverse, and many studies are single-case or developer-led, so results may not carry over to other learners or classrooms.
Frequently asked questions
How are single-case studies meta-analyzed?
With effect sizes suited to the data and multilevel models that treat observations as nested within cases and studies, alongside visual analysis and quality standards.
Why not use PND alone?
It is simple but sensitive to one extreme baseline point, has no sampling variance and shows ceiling effects, so better metrics are preferred.
Can studies with different disabilities be combined?
Only with care. We code how participants were identified and use disability type as a moderator.
Can placement studies show that inclusion works?
They are heavily confounded, so reviews are cautious and combine them with qualitative evidence.
Why does language matter?
Accurate, respectful terms avoid implying deficits as the cause of outcomes that are shaped by environments and supports.
Do you give advice about a child's education plan?
No. The service provides research and evidence-synthesis support only.
References
- Kratochwill TR, Hitchcock JH, Horner RH, et al. Single-case intervention research design standards. Remedial Spec Educ. 2013;34(1):26-38.
- Parker RI, Vannest KJ, Davis JL, Sauber SB. Combining nonoverlap and trend for single-case research: Tau-U. Behav Ther. 2011;42(2):284-299.
- Shadish WR, Hedges LV, Pustejovsky JE. Analysis and meta-analysis of single-case designs with a standardized mean difference statistic: a primer and applications. J Sch Psychol. 2014;52(2):123-147.
- Gersten R, Chard DJ, Jayanthi M, Baker SK, Morphy P, Flojo J. Mathematics instruction for students with learning disabilities: a meta-analysis of instructional components. Rev Educ Res. 2009;79(3):1202-1242.
- Steinbrenner JR, Hume K, Odom SL, et al. Evidence-based practices for children, youth, and young adults with autism. Chapel Hill: University of North Carolina, Frank Porter Graham Child Development Institute; 2020.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.