Meta-analysis and evidence synthesis for higher education research

Higher education research studies how students learn, persist and succeed at college and university, and how teaching, assessment, support and institutional policy affect them. Its evidence is a mix of course-level experiments, large administrative data sets, surveys and qualitative studies. Reviews have to handle selection into college, clustered data, outcomes measured differently across institutions and effects that depend on student groups.

Evidence synthesis in higher education

Higher education research covers teaching and learning at college and university, student retention and completion, assessment and feedback, advising and support, access and equity, and the effects of institutions and policy on student outcomes. Meta-analyses and systematic reviews have addressed topics such as active learning compared with lecturing, the effect of feedback, peer instruction, first-year programs, financial aid and the predictors of student persistence.

The evidence base has particular features. Students choose their institutions and courses, so comparisons across students or institutions are confounded by selection. Many experiments run within a single course and instructor. Outcomes such as grades and exam scores are not comparable across courses and institutions. Retention is a time-dependent outcome defined differently across systems. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for education.

Outcomes and their measurement

Common outcomes in higher education
OutcomeTypical measureIssue for synthesis
LearningExam scores, concept inventories, course gradesCourse-specific tests; grading rules and curves differ; instructor-written tests may favor the intervention
Persistence and retentionRe-enrolment next term or yearDefinitions differ (same institution versus any institution); time frame varies
CompletionDegree attainment within a time windowLong follow-up; transfer and part-time students complicate counting
EngagementSurvey scales (for example, national engagement surveys)Self-report; weak link with achievement
Graduate outcomesEmployment, earnings, further studyDepends on labor market and field; missing data

A review should state which outcomes it covers and keep them separate. Active learning studies, for example, report effects on exam scores and on pass rates, which are different outcomes and can be pooled only within type. Where studies use instructor-written tests aligned with the intervention, effects can be larger than with independent standardized tests, so test type is a useful moderator.

Selection and confounding

Students who take a course, join a program or attend an institution differ from those who do not. Studies that compare these groups without randomization may produce differences that reflect who chooses rather than what the program does. Better designs include randomized allocation, regression discontinuity based on admission or eligibility thresholds, difference-in-differences and matched comparisons with pre-program measures. A review should code design and baseline equivalence, such as prior achievement, and should interpret results from weaker designs cautiously.

Studies using institutional administrative data can have large samples that produce precise estimates, but precision does not remove confounding. A very precise estimate from a poorly controlled comparison is still biased, and a review should not give it more credibility than its design deserves.

Teaching methods and course-level evidence

Reviews of active learning in science, technology, engineering and mathematics courses have found higher exam scores and lower failure rates compared with traditional lecturing, on average, with variation across studies. Many of the included studies were quasi-experimental, and concerns were raised about whether instructors who adopt active methods differ in other ways. Subsequent work has looked at which components matter and for whom. Our reviews code the intensity of active learning, the instructor, the course level and the comparison condition, since "traditional lecturing" covers a variety of practices, and we show effects by design.

Other teaching topics include feedback, formative assessment, flipped classrooms, peer learning and technology use. Each has its own definition problems, and results depend on the quality of implementation, which is rarely measured well.

Equity, access and differential effects

Institutional programs may affect student groups differently, for example first-generation students, students from low-income families, students from under-represented groups or international students. Reviews should extract effects by group where studies report them, test group as a moderator and treat these analyses as exploratory unless planned. Subgroup results from single studies are often underpowered, and it is better to say that the evidence is limited than to infer that a program works for one group and not another. Language that treats student characteristics as causes of outcomes should be avoided, since the underlying mechanisms are usually institutional and social.

Retention and completion

Studies of retention model the probability of leaving over time as a function of student background, academic preparation, finances, integration and institutional features. Hazard models are common, and effect sizes are hazard ratios or odds ratios. Meta-analytic reviews of predictors, such as prior grades and motivation measures, have found that prior academic performance is among the stronger predictors. A review should be careful about interpretation: a predictor is not necessarily a lever for action. Intervention studies, such as advising, bridge programs and financial aid, are better suited to estimating what changes outcomes, and their effects are usually small to moderate and variable.

Institution, country and discipline

Institutional type (research university, community college, online provider), country system, discipline and student population all affect results. Most published experiments come from large research universities in the United States, often in introductory science courses. A review should present the distribution of settings and avoid claiming that findings apply to different systems. Meta-regression on institution or discipline is possible but suffers from small numbers and confounding. Comparing across countries is especially hard because the structure of degrees, admission and funding differ.

Publication bias and teaching innovations

Teaching innovations are often reported by the people who developed them, in settings where they are enthusiastic and well supported. Published evaluations may therefore overstate effects at scale. Failed innovations are less likely to be written up. We search conference proceedings and teaching journals, check for small-study effects and report developer-led and independent evaluations separately where possible.

Assessment, feedback and grading

Research on assessment and feedback covers formative assessment, rubrics, peer assessment, grading practices and the effects of testing on learning. Reviews of feedback have found large variation in effects, with feedback that addresses the task and offers a way forward typically helping more than praise or comparison with others. Studies differ in the type of feedback, its timing and who gave it, so a review needs a coding scheme for each, applied by two coders who work independently and compare results. Peer assessment studies raise their own questions: reliability of peer grades compared with staff grades, and effects on the reviewer as well as on the person reviewed.

Grading is not a stable measurement across institutions. Studies that use grade point average as an outcome mix different grading cultures, and a review should note that cross-institution grade comparisons rest on strong assumptions. Within-course comparisons using the same assessment are more credible than comparisons of grades across courses, and reviews should prefer them whenever they are available.

Coding and data extraction

The coding frame for a higher education review records the discipline, course level, class size, institution type, country, student population, instructor (whether the same instructor taught both conditions), assignment method, baseline measures, outcome measure and test type, follow-up time and attrition. For quasi-experiments, we record how groups were formed and which covariates were adjusted for. Two coders work on a sample and agreement is reported. Where a study reports results for several courses or terms, we decide in advance whether to treat them as one study with several estimates or as separate studies, and we justify the decision, because classes taught by one instructor are not independent. The coded data set and the decisions made are shared with the final report so that other researchers can reproduce the analysis and test alternative choices.

Common pitfalls we look for

  • Pooling instructor-written tests with standardized measures without a moderator.
  • Treating a comparison of self-selected groups as an intervention effect.
  • Counting multiple sections of one course as independent studies.
  • Mixing definitions of retention and completion.
  • Interpreting predictors of persistence as causes.
  • Over-reading subgroup differences from small samples.

Planning a higher education synthesis

We help define the intervention or predictor, the student population and the outcome, plan searches in ERIC, Scopus, Web of Science, PsycINFO and discipline-specific sources, and set up coding of design, baseline equivalence, test type, setting and implementation. For questions involving students' experience, we combine quantitative synthesis with qualitative evidence synthesis. See the systematic review service for scope and process.

An invented example of reading an effect

Suppose a review of an active-learning method finds an average standardized difference of 0.30 on exam scores across 40 invented studies, with a prediction interval from minus 0.10 to 0.70. The average is a small-to-moderate benefit. On a test with a standard deviation of 12 points, 0.30 corresponds to about 3.6 points. The prediction interval says that in some courses the effect could be negligible or slightly negative. If studies using independent standardized tests show an average of 0.18 and instructor-written tests 0.42, the instrument matters, and the more conservative figure might be the better basis for planning. These would be findings about courses already studied, and they do not tell a department what will happen in its own.

Ethical notes

Student data are sensitive. A review works from published, aggregated results and does not require student-level data. If a review plans to use institutional data directly, that requires separate approvals that are the responsibility of the institution and are outside this service.

How we support research projects in this area

Support

From a research question to a published synthesis

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A question, inclusion standards for study quality, and a plan for clustered designs and outcome types.

  • Searching and extraction

    Searches of education databases and repositories of evaluations, with double coding of effect sizes and moderators.

  • Synthesis

    Random-effects and robust variance models, adjustment for clustering, moderator analysis and sensitivity analysis.

  • Manuscript and submission

    The manuscript, coded-study tables and journal preparation.

Get a quoteDescribe the intervention, outcomes and target journal.

Boundaries of this service

A synthesis of higher education research describes average effects across studies and settings. It does not provide institutional policy, admissions or advising decisions, or evaluation of any student, course or institution. Many studies come from a few countries and course types, and associations with student outcomes are not necessarily causal.

Frequently asked questions

Does active learning improve outcomes in college courses?

On average, reviews find higher exam scores and lower failure rates than lecturing, with variation by design, course and implementation.

How is student selection handled?

By coding design and baseline equivalence, favoring randomized and strong quasi-experimental studies, and interpreting weaker designs with caution.

Why separate instructor-written from standardized tests?

Tests aligned with the intervention can show larger effects, so test type is examined as a moderator.

Can predictors of retention be treated as causes?

No. They describe who is more likely to persist. Intervention studies are needed to see what changes outcomes.

Do findings apply to all institutions?

Not necessarily. Most studies come from research universities in a few countries, and the review reports the coverage.

Do you give institutional policy advice?

No. The service provides research and evidence-synthesis support only.

References

  1. Freeman S, Eddy SL, McDonough M, et al. Active learning increases student performance in science, engineering, and mathematics. Proc Natl Acad Sci U S A. 2014;111(23):8410-8415.
  2. Hattie J, Timperley H. The power of feedback. Rev Educ Res. 2007;77(1):81-112.
  3. Robbins SB, Lauver K, Le H, Davis D, Langley R, Carlstrom A. Do psychosocial and study skill factors predict college outcomes? A meta-analysis. Psychol Bull. 2004;130(2):261-288.
  4. Sneyers E, De Witte K. Interventions in higher education and their effect on student success: a meta-analysis. Educ Rev. 2018;70(2):208-228.
  5. Theobald EJ, Hill MJ, Tran E, et al. Active learning narrows achievement gaps for underrepresented students in undergraduate science, technology, engineering, and math. Proc Natl Acad Sci U S A. 2020;117(12):6476-6483.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.