Evidence synthesis in higher education
Higher education research covers teaching and learning at college and university, student retention and completion, assessment and feedback, advising and support, access and equity, and the effects of institutions and policy on student outcomes. Meta-analyses and systematic reviews have addressed topics such as active learning compared with lecturing, the effect of feedback, peer instruction, first-year programs, financial aid and the predictors of student persistence.
The evidence base has particular features. Students choose their institutions and courses, so comparisons across students or institutions are confounded by selection. Many experiments run within a single course and instructor. Outcomes such as grades and exam scores are not comparable across courses and institutions. Retention is a time-dependent outcome defined differently across systems. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for education.
Outcomes and their measurement
| Outcome | Typical measure | Issue for synthesis |
|---|---|---|
| Learning | Exam scores, concept inventories, course grades | Course-specific tests; grading rules and curves differ; instructor-written tests may favor the intervention |
| Persistence and retention | Re-enrolment next term or year | Definitions differ (same institution versus any institution); time frame varies |
| Completion | Degree attainment within a time window | Long follow-up; transfer and part-time students complicate counting |
| Engagement | Survey scales (for example, national engagement surveys) | Self-report; weak link with achievement |
| Graduate outcomes | Employment, earnings, further study | Depends on labor market and field; missing data |
A review should state which outcomes it covers and keep them separate. Active learning studies, for example, report effects on exam scores and on pass rates, which are different outcomes and can be pooled only within type. Where studies use instructor-written tests aligned with the intervention, effects can be larger than with independent standardized tests, so test type is a useful moderator.
Selection and confounding
Students who take a course, join a program or attend an institution differ from those who do not. Studies that compare these groups without randomization may produce differences that reflect who chooses rather than what the program does. Better designs include randomized allocation, regression discontinuity based on admission or eligibility thresholds, difference-in-differences and matched comparisons with pre-program measures. A review should code design and baseline equivalence, such as prior achievement, and should interpret results from weaker designs cautiously.
Studies using institutional administrative data can have large samples that produce precise estimates, but precision does not remove confounding. A very precise estimate from a poorly controlled comparison is still biased, and a review should not give it more credibility than its design deserves.
Teaching methods and course-level evidence
Reviews of active learning in science, technology, engineering and mathematics courses have found higher exam scores and lower failure rates compared with traditional lecturing, on average, with variation across studies. Many of the included studies were quasi-experimental, and concerns were raised about whether instructors who adopt active methods differ in other ways. Subsequent work has looked at which components matter and for whom. Our reviews code the intensity of active learning, the instructor, the course level and the comparison condition, since "traditional lecturing" covers a variety of practices, and we show effects by design.
Other teaching topics include feedback, formative assessment, flipped classrooms, peer learning and technology use. Each has its own definition problems, and results depend on the quality of implementation, which is rarely measured well.
Equity, access and differential effects
Institutional programs may affect student groups differently, for example first-generation students, students from low-income families, students from under-represented groups or international students. Reviews should extract effects by group where studies report them, test group as a moderator and treat these analyses as exploratory unless planned. Subgroup results from single studies are often underpowered, and it is better to say that the evidence is limited than to infer that a program works for one group and not another. Language that treats student characteristics as causes of outcomes should be avoided, since the underlying mechanisms are usually institutional and social.
Retention and completion
Studies of retention model the probability of leaving over time as a function of student background, academic preparation, finances, integration and institutional features. Hazard models are common, and effect sizes are hazard ratios or odds ratios. Meta-analytic reviews of predictors, such as prior grades and motivation measures, have found that prior academic performance is among the stronger predictors. A review should be careful about interpretation: a predictor is not necessarily a lever for action. Intervention studies, such as advising, bridge programs and financial aid, are better suited to estimating what changes outcomes, and their effects are usually small to moderate and variable.
Institution, country and discipline
Institutional type (research university, community college, online provider), country system, discipline and student population all affect results. Most published experiments come from large research universities in the United States, often in introductory science courses. A review should present the distribution of settings and avoid claiming that findings apply to different systems. Meta-regression on institution or discipline is possible but suffers from small numbers and confounding. Comparing across countries is especially hard because the structure of degrees, admission and funding differ.
Publication bias and teaching innovations
Teaching innovations are often reported by the people who developed them, in settings where they are enthusiastic and well supported. Published evaluations may therefore overstate effects at scale. Failed innovations are less likely to be written up. We search conference proceedings and teaching journals, check for small-study effects and report developer-led and independent evaluations separately where possible.
Assessment, feedback and grading
Research on assessment and feedback covers formative assessment, rubrics, peer assessment, grading practices and the effects of testing on learning. Reviews of feedback have found large variation in effects, with feedback that addresses the task and offers a way forward typically helping more than praise or comparison with others. Studies differ in the type of feedback, its timing and who gave it, so a review needs a coding scheme for each, applied by two coders who work independently and compare results. Peer assessment studies raise their own questions: reliability of peer grades compared with staff grades, and effects on the reviewer as well as on the person reviewed.
Grading is not a stable measurement across institutions. Studies that use grade point average as an outcome mix different grading cultures, and a review should note that cross-institution grade comparisons rest on strong assumptions. Within-course comparisons using the same assessment are more credible than comparisons of grades across courses, and reviews should prefer them whenever they are available.
Coding and data extraction
The coding frame for a higher education review records the discipline, course level, class size, institution type, country, student population, instructor (whether the same instructor taught both conditions), assignment method, baseline measures, outcome measure and test type, follow-up time and attrition. For quasi-experiments, we record how groups were formed and which covariates were adjusted for. Two coders work on a sample and agreement is reported. Where a study reports results for several courses or terms, we decide in advance whether to treat them as one study with several estimates or as separate studies, and we justify the decision, because classes taught by one instructor are not independent. The coded data set and the decisions made are shared with the final report so that other researchers can reproduce the analysis and test alternative choices.
Common pitfalls we look for
- Pooling instructor-written tests with standardized measures without a moderator.
- Treating a comparison of self-selected groups as an intervention effect.
- Counting multiple sections of one course as independent studies.
- Mixing definitions of retention and completion.
- Interpreting predictors of persistence as causes.
- Over-reading subgroup differences from small samples.
Planning a higher education synthesis
We help define the intervention or predictor, the student population and the outcome, plan searches in ERIC, Scopus, Web of Science, PsycINFO and discipline-specific sources, and set up coding of design, baseline equivalence, test type, setting and implementation. For questions involving students' experience, we combine quantitative synthesis with qualitative evidence synthesis. See the systematic review service for scope and process.
An invented example of reading an effect
Suppose a review of an active-learning method finds an average standardized difference of 0.30 on exam scores across 40 invented studies, with a prediction interval from minus 0.10 to 0.70. The average is a small-to-moderate benefit. On a test with a standard deviation of 12 points, 0.30 corresponds to about 3.6 points. The prediction interval says that in some courses the effect could be negligible or slightly negative. If studies using independent standardized tests show an average of 0.18 and instructor-written tests 0.42, the instrument matters, and the more conservative figure might be the better basis for planning. These would be findings about courses already studied, and they do not tell a department what will happen in its own.
Ethical notes
Student data are sensitive. A review works from published, aggregated results and does not require student-level data. If a review plans to use institutional data directly, that requires separate approvals that are the responsibility of the institution and are outside this service.
How we support research projects in this area
From a research question to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A question, inclusion standards for study quality, and a plan for clustered designs and outcome types.
Searching and extraction
Searches of education databases and repositories of evaluations, with double coding of effect sizes and moderators.
Synthesis
Random-effects and robust variance models, adjustment for clustering, moderator analysis and sensitivity analysis.
Manuscript and submission
The manuscript, coded-study tables and journal preparation.
Boundaries of this service
A synthesis of higher education research describes average effects across studies and settings. It does not provide institutional policy, admissions or advising decisions, or evaluation of any student, course or institution. Many studies come from a few countries and course types, and associations with student outcomes are not necessarily causal.
Frequently asked questions
Does active learning improve outcomes in college courses?
On average, reviews find higher exam scores and lower failure rates than lecturing, with variation by design, course and implementation.
How is student selection handled?
By coding design and baseline equivalence, favoring randomized and strong quasi-experimental studies, and interpreting weaker designs with caution.
Why separate instructor-written from standardized tests?
Tests aligned with the intervention can show larger effects, so test type is examined as a moderator.
Can predictors of retention be treated as causes?
No. They describe who is more likely to persist. Intervention studies are needed to see what changes outcomes.
Do findings apply to all institutions?
Not necessarily. Most studies come from research universities in a few countries, and the review reports the coverage.
Do you give institutional policy advice?
No. The service provides research and evidence-synthesis support only.
References
- Freeman S, Eddy SL, McDonough M, et al. Active learning increases student performance in science, engineering, and mathematics. Proc Natl Acad Sci U S A. 2014;111(23):8410-8415.
- Hattie J, Timperley H. The power of feedback. Rev Educ Res. 2007;77(1):81-112.
- Robbins SB, Lauver K, Le H, Davis D, Langley R, Carlstrom A. Do psychosocial and study skill factors predict college outcomes? A meta-analysis. Psychol Bull. 2004;130(2):261-288.
- Sneyers E, De Witte K. Interventions in higher education and their effect on student success: a meta-analysis. Educ Rev. 2018;70(2):208-228.
- Theobald EJ, Hill MJ, Tran E, et al. Active learning narrows achievement gaps for underrepresented students in undergraduate science, technology, engineering, and math. Proc Natl Acad Sci U S A. 2020;117(12):6476-6483.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.