Evidence synthesis in e-learning
E-learning research asks whether and how learning with digital technology works. Questions include how online courses compare with face-to-face ones, whether blended designs outperform either, what role interaction and feedback play, why many learners drop out of open online courses, and whether adaptive or analytics-driven systems help. A widely cited review by the US Department of Education found that blended learning outperformed face-to-face teaching on average, and that purely online and face-to-face were similar, but that the advantage of blended learning was linked to extra time and instructional elements, not the medium itself.
That point is central to the field. Technology is a delivery channel, and the effect of a study depends on the teaching method it carries. Reviews need to hold constant what is taught and how when they compare media. They also face fast-changing tools, high dropout, self-selection into online courses and many studies of low quality. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for education.
Medium versus method
| Comparison | What differs | Issue for synthesis |
|---|---|---|
| Online vs face-to-face, same content and teacher | Delivery mode | Closest to a test of the medium, but rare |
| Online vs face-to-face, different courses | Mode plus content, instruction and students | Confounded; effect cannot be attributed to the medium |
| Blended vs face-to-face | Mode plus extra time and materials | Advantage may reflect added learning time |
| Feature present vs absent within online | A design feature such as feedback or quizzes | Tests a method; strongest evidence for design |
| Technology vs no technology | Presence of a tool | Depends on the alternative; positive effects are common |
A review should classify each comparison and not pool across types without a moderator. Within-online comparisons of design features, such as practice quizzes, worked examples or feedback, give more useful information to course designers than overall comparisons of modes. Effect sizes against a weak control, such as a lecture with no active elements, are larger than against a strong alternative, and the comparator should be coded.
Dropout, completion and self-selection
Massive open online courses have completion rates that are often low, in the range of a small share of those who enrol, and many enrolees never start. Completion depends on intent, since many enrol to browse. A review should note how each study defined enrolment, activity and completion, and should avoid comparing completion rates across courses with different definitions. Studies of students in online degree programs face self-selection: those who choose online study differ from those who choose campus study. Where randomization is not possible, adjustment for prior achievement and other covariates helps but does not remove the problem.
Differential dropout also threatens comparisons of outcomes. If weaker learners drop out of one arm, the remaining learners score higher, which looks like a benefit. We extract attrition by arm and check whether results are sensitive to it.
Learning analytics and prediction
Learning analytics uses log data to predict who is at risk and to adapt instruction. Studies report predictive accuracy, such as area under the curve or classification accuracy, and sometimes the effect of an intervention triggered by a prediction. Predictive accuracy is not the same as educational benefit. A model that flags students correctly is useful only if the follow-up helps them. A review should separate prediction studies from intervention studies, report how models were validated (within a course or across courses and years) and note that models trained on one course often perform worse on another. Reporting follows the standards used for prediction models, which our prognostic meta-analysis page describes.
Interaction, feedback and design features
Meta-analyses of online and distance education have examined the roles of student-student, student-teacher and student-content interaction, and found that designs with more interaction tend to show better outcomes. Interaction is often measured as a design feature rather than as what learners did, and its quality is hard to code. Reviews can code features such as synchronous contact, feedback, collaborative tasks and the use of multimedia, and test them as moderators. The evidence for the multimedia principles (for instance, learning from words and pictures) comes largely from short laboratory experiments with students, and its extension to full courses needs care, because attention and workload over a whole course differ from those in a twenty-minute experiment.
Mobile, adaptive and game-based learning
Studies of mobile learning, intelligent tutoring, adaptive practice and educational games report positive average effects with wide variation. Adaptive systems in mathematics have performed well in some large trials and not in others. Game-based studies are often short and compare a game with a passive control, which favors the game. A review should code duration, control condition, learner age and subject, and should not combine games designed for fun with systems designed as drills. Newer tools, including those that generate text with language models, are developing faster than the research; reviews should state the period of the studies and the tools they covered.
Learners, context and access
Online learning requires devices, connectivity and skills in self-regulation. Studies done with motivated adult students in well-resourced settings may not apply to younger learners, learners with limited access or those who need more support. Reviews should report learner age, prior experience and country, and be careful with claims of equal benefit. Evidence on accessibility for learners with disabilities exists in small studies and in guidelines, and a review might map it rather than pool it.
Publication bias and rapid obsolescence
Studies of new tools are often published by the people who built them, in favorable settings, and show positive results. Small studies with strong effects are more common than large null studies. We assess small-study effects and compare developer-led with independent evaluations. Studies of older technologies may no longer be relevant, so the date of data collection is coded for every study and the review discusses whether the conclusions are specific to a period.
Engagement measures and log data
Online platforms record clicks, views, time on page and submissions, and studies use these as measures of engagement. They are convenient and also crude: time on a page does not show attention, and a learner may leave a tab open while doing something else. Correlations between log measures and learning are typically modest and differ across courses. A review should record which log variables were used, how they were defined, and whether engagement was treated as an outcome, a mediator or a predictor. Surveys of engagement and satisfaction are common in this field and have their own limits as self-reports; they are reported separately from learning outcomes, not combined with them, because learners can enjoy a course without learning more from it and can learn a good deal from a course they do not enjoy.
Qualitative and survey evidence
Much e-learning research is descriptive: interviews, focus groups and surveys on learners' and teachers' experiences. These studies give information on barriers such as isolation, workload, technical problems and the need for self-discipline, and on what teachers found hard when moving to online delivery. They are not pooled numerically. A qualitative evidence synthesis can bring them together, using methods such as thematic synthesis, and can be linked to quantitative results to suggest why effects differ across settings. Survey data on satisfaction are often collected from learners who completed a course, which leaves out those who left, and reviews should state this limitation when summarizing satisfaction.
Common pitfalls we look for
- Attributing differences to the medium when content and method also differ.
- Comparing completion rates with different definitions of enrolment.
- Ignoring differential dropout between arms.
- Treating prediction accuracy as educational benefit.
- Pooling games, drills and simulations as one intervention.
- Applying results for one learner population to all.
Planning an e-learning synthesis
We help define the technology, the instructional method, the learners and the outcome, plan searches in ERIC, Scopus, Web of Science, IEEE Xplore and the ACM Digital Library, and set up coding of comparison type, design features, duration, attrition and period. See the systematic review service for scope and process.
An invented example of reading a mode comparison
Suppose a review reports that blended courses outperform face-to-face courses by 0.25 standard deviations across 30 invented studies, but that in the 8 studies where instructors and content were the same the difference is 0.05 and in the 10 where blended learners had extra time it is 0.40. The headline number is then a mixture. The medium itself seems to add little, and the larger effects come from added time and different instruction. A course planner would learn that the design choices matter more than whether the course is online, a message that a single pooled figure would conceal.
Coding and transparency
Coding records the mode, platform type, synchronous or asynchronous contact, instructor presence, course length, subject, learner level, comparison condition, assignment method, attrition by arm, outcome type and study period. Two coders work independently on a sample, agreement is reported and the coded data and code are shared with the final report. When a study describes the intervention only briefly, we contact authors and record whether we received details.
How we support research projects in this area
From a research question to a published synthesis
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A question, inclusion standards for study quality, and a plan for clustered designs and outcome types.
Searching and extraction
Searches of education databases and repositories of evaluations, with double coding of effect sizes and moderators.
Synthesis
Random-effects and robust variance models, adjustment for clustering, moderator analysis and sensitivity analysis.
Manuscript and submission
The manuscript, coded-study tables and journal preparation.
Boundaries of this service
An e-learning synthesis describes average effects of digital learning designs across studies. It does not select platforms, design courses or evaluate any learning product or provider. Technology changes faster than research, so evidence may not describe current tools, and effects depend on the teaching method, learners and context.
Frequently asked questions
Is online learning as effective as face-to-face?
On average, reviews find similar results, with advantages for blended designs that often reflect added time or different instruction. The medium alone adds little.
Why are completion rates hard to compare?
Studies define enrolment, activity and completion differently, and many enrolees never intend to finish.
Does a learning analytics model improve learning?
Not by itself. Accuracy in predicting risk is not the same as benefit; intervention studies are needed.
Can game-based learning results be pooled?
Only within comparable types, with duration and control condition coded, since games differ greatly.
Do results from adult online courses apply to children?
Not necessarily. Age, self-regulation and access differ, so learner population is reported and tested.
Do you recommend platforms or course designs?
No. The service provides research and evidence-synthesis support only.
References
- Means B, Toyama Y, Murphy R, Bakia M, Jones K. Evaluation of evidence-based practices in online learning: a meta-analysis and review of online learning studies. Washington, DC: US Department of Education; 2010.
- Bernard RM, Abrami PC, Borokhovski E, et al. A meta-analysis of three types of interaction treatments in distance education. Rev Educ Res. 2009;79(3):1243-1289.
- Jordan K. Initial trends in enrolment and completion of massive open online courses. Int Rev Res Open Distrib Learn. 2014;15(1):133-160.
- Clark RC, Mayer RE. E-learning and the science of instruction. 4th ed. Hoboken: Wiley; 2016.
- Kulik JA, Fletcher JD. Effectiveness of intelligent tutoring systems: a meta-analytic review. Rev Educ Res. 2016;86(1):42-78.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.