Meta-analysis and evidence synthesis for STEM education research

STEM education research studies how students learn science, technology, engineering and mathematics, and which teaching approaches, curricula and tools help. Its evidence ranges from classroom experiments to national assessments. Reviews have to handle outcome measures that vary in what they test, small and short studies, and the gap between laboratory-style trials and real classrooms.

Evidence synthesis in STEM education

STEM education research asks how to teach science, mathematics, engineering and computing so that students learn, persist and see themselves as capable. Common topics include inquiry-based and active learning, problem-based learning, mathematics instruction and intervention for struggling learners, educational technology and simulations, computing and programming education, and the participation of women and under-represented groups. Meta-analyses have addressed many of these, with average effects usually small to moderate and variable.

The evidence has features that need care. Interventions are complex, with several components that vary between studies. Outcome measures range from short tests written by researchers to standardized assessments. Many studies are small, last a few weeks, and are conducted by developers. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for education.

Outcome measures and what they test

Types of STEM outcome measures
Measure typeExampleIssue for synthesis
Researcher-designed testTest written for the studyAligned with the intervention; effects often larger than on independent tests
Standardized assessmentNational or commercial testLess sensitive to intervention content; effects tend to be smaller
Concept inventoryInstrument assessing common misconceptionsPre-post gain scores; normalized gain is not an effect size in the usual sense
Problem-solving taskOpen-ended or transfer problemsScoring rubrics vary; reliability not always reported
Attitude or self-efficacy scaleSurvey of interest and confidenceSelf-report; weakly related to later choices

Systematic differences in effect between researcher-made and standardized measures are well documented in education. A review should code the measure type and report results separately or test type as a moderator. For concept inventories, the normalized gain (the share of possible improvement achieved) is popular in physics education research, but it is not a standardized mean difference, and its variance properties are not simple. We either convert to standard effect sizes using the reported means and standard deviations or analyze gains separately, and we explain the choice.

Complex interventions and what counts as the same

"Inquiry-based learning" and "active learning" describe families of approaches, not single treatments. They differ in the amount of guidance students receive, the use of technology, the time spent, the role of teachers and the degree of collaboration. Evidence suggests that unguided discovery learning tends to be less effective than guided forms, so pooling all inquiry studies together can hide that difference. A review should define the intervention components in the protocol, code each study for them and analyze components as moderators.

Fidelity of implementation is rarely reported well. Intended effects may not occur when teachers lack training or time. Where studies report fidelity or teacher training, we use those variables in moderator analysis.

Design, duration and developer effects

Many STEM education studies are quasi-experimental with intact classes, so groups may differ at the start. Short interventions lasting a few lessons may produce effects that do not persist or scale. Studies run by those who developed the intervention tend to show larger effects than independent evaluations. A review should code design (randomized, matched, pre-post only), baseline equivalence, duration, developer involvement and follow-up, and compare effects by each. Results from single-group pre-post studies should not be pooled with controlled studies, because a gain over time occurs without any intervention.

Classes and schools are the unit of assignment in many trials, so analyses that treat students as independent exaggerate precision. Reviews should extract the cluster information and adjust or use the cluster-adjusted effect sizes.

Mathematics learning and interventions

Research on mathematics includes core instruction, interventions for students with difficulties, use of manipulatives and representations, and the teaching of problem solving. Meta-analyses of interventions for struggling learners have found that explicit, systematic instruction produces positive effects on average, and that effects are larger for researcher-designed measures. Studies often involve small groups taught by researchers, so effects at scale in ordinary classrooms may be smaller. A review should describe the group size, who delivered the intervention and the measures used, and should not generalize beyond these conditions.

Computing and engineering education

Computing education has grown quickly, and reviews cover programming instruction, pair programming, block-based environments and the effects of computational thinking activities. Evidence is more recent and more varied, with many small studies, different outcome measures and little standardization of tests. Engineering education studies frequently describe project-based courses in single institutions. A review should be honest about the limited number of controlled studies and use narrative synthesis where pooling is not justified.

Gender, background and participation

Research on participation looks at gender gaps, effects on under-represented groups and the role of beliefs about ability. Studies of gender differences in mathematics achievement have found small average differences that vary across countries and measures, and larger differences in participation in some fields. Meta-analysis of differences between groups describes patterns and should not be taken to imply innate causes. Intervention studies that aim to narrow gaps, such as value-affirmation exercises, have shown mixed results in replications, and a review should report how effects differ across settings and whether the original findings replicated.

Publication bias and replication

Small studies by program developers with large effects are the most likely to be published. We check small-study effects, compare published and unpublished work and weight conclusions toward studies with better designs. Large replication trials of promising interventions have found smaller effects than early small studies, which is a reason to look at the pattern across study size and independence, not only the pooled mean.

Educational technology, simulations and laboratories

Studies of simulations, virtual laboratories, intelligent tutoring systems and games in STEM report positive average effects, with considerable variation. What distinguishes effective from ineffective uses is often the instructional design around the technology, such as whether students receive guidance and feedback, rather than the technology itself. A review should code the type of technology, the role it plays (replacing or supplementing instruction), the comparison condition and the duration. Comparisons of technology with no instruction differ from comparisons with well-designed conventional teaching, and effects are larger for the former. Where studies compare a laboratory with a simulation of the same activity, effects tend to be small, which tells a different story from the larger gains against a weak control.

Technology changes quickly, so older studies may describe tools no longer in use. Period of study is coded and tested as a moderator, and conclusions about specific products are avoided.

Teachers, professional development and scale

The effect of a curriculum or method depends on teachers, and many studies of professional development report changes in teacher knowledge or practice and, less often, in student outcomes. Reviews of professional development have found that programs with more contact hours and a focus on subject content and student thinking tend to show stronger effects, although associations between duration and effects are not consistent. A review should separate outcomes for teachers from outcomes for students and describe how many studies report the second. When effects are measured at the class level, few classes per study mean wide intervals.

Scaling up an intervention changes who delivers it, the support available and the context. Studies at scale in many schools are valuable because they test performance under realistic conditions, and a review should highlight them and compare with small trials.

Common pitfalls we look for

  • Pooling researcher-made and standardized tests without a moderator.
  • Treating all "inquiry" or "active learning" as one intervention.
  • Using single-group pre-post gains as intervention effects.
  • Ignoring clustering by class or school.
  • Overreading short-term effects.
  • Interpreting group differences as explanations of the differences.

Planning a STEM education synthesis

We help define the intervention and its components, the learner group and the outcome, plan searches in ERIC, Scopus, Web of Science and discipline-based education research sources, and set up coding of design, measure type, duration, developer involvement and setting. Where studies are too varied to pool, we plan a structured narrative synthesis. See the systematic review service for scope and process.

An invented example of interpreting a small effect

Suppose a review finds an average standardized difference of 0.20 for a science teaching approach on independent standardized tests across 25 invented studies, and 0.45 on researcher-made tests. On a test with a standard deviation of 15 points, 0.20 is about 3 points. A reader would reasonably ask whether a gain of this size justifies the cost of training teachers, and the answer depends on local conditions the review cannot see. The more informative message is the contrast between the two measure types, which suggests that part of the apparent benefit reflects alignment of the test to the teaching, and that planners should expect something closer to the smaller figure in settings that use independent assessments.

Coding and transparency

Coding frames record intervention components, dosage (hours and weeks), teacher training, grade level, subject, country, class size, assignment method, baseline equivalence, measure type and developer involvement. Two coders work on a sample, agreement is reported and the coded data and code are shared with the report. When effects are computed from means and standard deviations, we record the source of each value, and when conversions are needed from other statistics we state the formula used.

How we support research projects in this area

Support

From a research question to a published synthesis

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A question, inclusion standards for study quality, and a plan for clustered designs and outcome types.

  • Searching and extraction

    Searches of education databases and repositories of evaluations, with double coding of effect sizes and moderators.

  • Synthesis

    Random-effects and robust variance models, adjustment for clustering, moderator analysis and sensitivity analysis.

  • Manuscript and submission

    The manuscript, coded-study tables and journal preparation.

Get a quoteDescribe the intervention, outcomes and target journal.

Boundaries of this service

A STEM education synthesis describes average effects across studies and settings. It does not provide curriculum design, teacher evaluation or recommendations for any school, class or student. Many studies are short and conducted by developers under favorable conditions, so effects in everyday practice may be different.

Frequently asked questions

Is inquiry-based learning effective?

Reviews find positive effects on average, with guided forms generally doing better than unguided discovery. A review codes components to show this.

Why do effects differ between test types?

Researcher-made tests align with the intervention and show larger effects than independent standardized tests, so type is a moderator.

How do you treat concept inventory gains?

By converting to standard effect sizes when means and standard deviations are available, or by analyzing normalized gains separately with an explanation.

Do small developer-led studies overstate effects?

They can. We code developer involvement and compare with independent and larger studies.

What about gender gap research?

We describe patterns across studies and settings and avoid implying causes that the data cannot show.

Do you provide curriculum design or teacher evaluation?

No. The service provides research and evidence-synthesis support only.

References

  1. Kirschner PA, Sweller J, Clark RE. Why minimal guidance during instruction does not work: an analysis of the failure of constructivist, discovery, problem-based, experiential, and inquiry-based teaching. Educ Psychol. 2006;41(2):75-86.
  2. Furtak EM, Seidel T, Iverson H, Briggs DC. Experimental and quasi-experimental studies of inquiry-based science teaching: a meta-analysis. Rev Educ Res. 2012;82(3):300-329.
  3. Gersten R, Chard DJ, Jayanthi M, Baker SK, Morphy P, Flojo J. Mathematics instruction for students with learning disabilities: a meta-analysis of instructional components. Rev Educ Res. 2009;79(3):1202-1242.
  4. Cheung AC, Slavin RE. How methodological features affect effect sizes in education. Educ Res. 2016;45(5):283-292.
  5. Hyde JS, Lindberg SM, Linn MC, Ellis AB, Williams CC. Gender similarities characterize math performance. Science. 2008;321(5888):494-495.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.