Evidence synthesis in computer science and AI
Computer science studies computation, software and information systems; artificial intelligence studies methods that let machines perform tasks that normally involve human judgment, such as perception and language. Evidence syntheses in the field take different forms. Systematic literature reviews and mapping studies are standard in empirical software engineering, following guidelines adapted from medicine. Meta-analyses of controlled experiments exist on topics such as pair programming, test-driven development and the effects of interface design. Benchmark comparisons in machine learning are summarized in survey papers and leaderboards that do not follow systematic-review rules.
The features of this evidence call for care. Human-subject experiments are small and often use students. Software projects differ in size, domain and process. Machine learning results depend on data splits, tuning budgets and random seeds, and benchmark data may leak into training sets. Technology changes quickly, which dates findings. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for engineering and technology. The service does not build software or models.
Types of evidence and how they are synthesized
| Type | Example | How it can be synthesized |
|---|---|---|
| Controlled experiment with people | Developers using two tools on a task | Meta-analysis of standardized differences; moderators for participants and task |
| Repository mining | Defect rates across open-source projects | Descriptive synthesis or meta-analysis of correlations; confounded by project size |
| Survey of practitioners | Reported use of practices | Prevalence synthesis; sampling bias |
| Benchmark comparison | Accuracy of models on a data set | Tabulation; careful pooling only for comparable settings; variance from seeds |
| Qualitative case study | Adoption of a process in a company | Thematic synthesis; mapping |
Because the evidence types differ, one review seldom uses all of them. A systematic mapping study is a good first step: it classifies studies by topic, method and context and shows where evidence is concentrated, without trying to estimate effects. Full meta-analysis becomes possible when there are enough comparable experiments with reported means, standard deviations and sample sizes. The scoping review page describes mapping approaches.
Experiments with developers and users
Software engineering experiments often compare a practice or tool in a task lasting hours, with 20 to 60 participants, many of them students. Results may not transfer to professionals working on large systems over months. Reviews should code participant experience, task size and realism, the comparator and whether outcomes were measured objectively (time, defects) or subjectively (perceived ease). Studies that report only p-values without effect sizes complicate synthesis, and the review should try to compute effects from reported statistics or ask authors. Meta-analyses in this area typically find small to moderate average effects with large heterogeneity, which reflects genuine variation in tasks and teams.
Human-computer interaction user studies add measures such as usability scales and task time. Within-subject designs with order effects need the correct handling in synthesis, as discussed for crossover designs elsewhere on this site.
Machine learning benchmarks and comparisons
Papers in machine learning report improvements on benchmarks, usually as accuracy or a similar score, often as a single run. Differences of one or two points can be within the variation from random seeds, data splits and hyperparameter search. Reviews of this literature have documented that baselines are sometimes weakly tuned compared to the proposed method, that test sets are reused many times, and that benchmark results do not always predict performance on new data. A synthesis should record the data set and split, tuning budget, number of runs and variance, and whether code and data were released. It should avoid ranking methods across papers that used different protocols. Where a fair comparison exists, a controlled re-evaluation by independent researchers is better evidence than numbers copied from original papers.
Data leakage, reproducibility and reporting
Leakage occurs when information from the test set influences training, for example through shared samples, preprocessing on the full data set or duplicates between train and test partitions. It inflates performance and is surprisingly common, including in applied fields that use machine learning. Reproducibility depends on code, data, environment and seeds, and checklists have been proposed to improve reporting. A review should assess whether primary studies guarded against leakage and reported enough to be reproduced, and should report the proportion that did. Studies of prediction models in medicine using machine learning have their own risk-of-bias tools, discussed on the prognostic meta-analysis page.
Fairness, safety and evaluation of AI systems
Research on the fairness, robustness and safety of AI systems uses audits, benchmarks and user studies. Measures of fairness differ and can conflict, and results depend on the data and the population. A review should describe which definitions were used, how groups were defined and what the evaluation data represented. Evaluations of large language models are particularly sensitive to prompt wording, contamination of test data and rapid model changes, so findings may describe a model version that is no longer in use. Such studies should be dated and linked to versions, and claims about general capability should be avoided. This service does not evaluate or certify any system.
Rapid change and shelf life
Computing evidence ages quickly. A study of a programming practice from 2005 may not reflect current tools, and a machine learning result from two years ago may concern models far less capable than the present ones. Reviews should state the period of the evidence, use publication year or technology generation as a moderator, and give a date after which findings should be treated cautiously. Living reviews help where a topic moves fast and where decisions depend on current evidence.
Publication bias and selective reporting
The field rewards new methods that beat baselines, and null or negative results are seldom published. In experimental software engineering, publication bias tests have found evidence of small-study effects. In machine learning, only methods that improve a benchmark appear in papers, and failed approaches go unreported. We search preprints, theses and conference papers, which are the main venues in computing, and report differences between sources.
Process, productivity and practitioner evidence
Studies of software process, agile methods, code review, continuous integration and developer productivity use surveys, repository data and case studies. Repository studies can analyze thousands of projects, but projects differ in size, age and purpose, and measures such as lines of code or commit counts are poor proxies for productivity or quality. Associations between a practice and defects are confounded by project size and team experience, and many studies do not adjust for them. A review should record the measures, the controls and the project selection, and treat large samples as a source of precision, not of validity. Surveys of practitioners are affected by who responds, and self-reported use of practices tends to overstate adoption.
Qualitative studies of teams and organizations explain why practices work in some places and not others. They can be combined by thematic synthesis and linked to quantitative findings in a mixed-methods review, which helps readers see the conditions under which a practice appears to help.
Common pitfalls we look for
- Treating student experiments as evidence about professional practice.
- Ranking methods across papers with different protocols.
- Ignoring variance from seeds and tuning.
- Missing data leakage between training and test sets.
- Reporting p-values without effect sizes.
- Presenting findings about one model version as general.
Planning a computer science or AI synthesis
We help define the question and choose between a mapping study, a meta-analysis of experiments and a structured comparison of benchmark results, plan searches in IEEE Xplore, ACM Digital Library, Scopus, Web of Science, arXiv and DBLP, and set up coding of participants, task, tool, data set, split, tuning, runs, variance, code availability and dates. See the systematic review service for scope and process.
An invented example of benchmark noise
Suppose a new method reports 91.2 percent accuracy and a baseline 90.5 percent on the same test set of 2,000 items. The difference of 0.7 points is 14 items. A rough standard error of a single accuracy near 91 percent on 2,000 items is the square root of 0.91 times 0.09 divided by 2,000, about 0.64 points, so a difference of 0.7 points is of the same order as the sampling error of one score. Because both methods are tested on the same items, a paired test has a smaller error than this, which is why paired tests such as McNemar's test are preferred, but variation from seeds and tuning comes on top. A review that counted this paper as a win for the new method would be counting noise. It would instead record the number of runs and variation, and ask whether independent evaluations confirm the gain.
Coding and transparency
Coding frames record study type, participants and experience, task and realism, tools or methods compared, outcome and how it was measured, data sets and splits, tuning budgets, number of runs, variance, availability of code and data, and dates. Two coders work independently on a sample, and the coded data and scripts are shared with the final report.
How we support research projects in this area
From a technical question to a published review
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
Help choosing between a mapping study, a systematic review and a meta-analysis, and a protocol with research questions.
Searching and extraction
Database search, snowballing and grey literature, with screening and a classification scheme.
Synthesis
Structured synthesis, quality assessment and, where experiments are comparable, meta-analysis.
Manuscript and submission
The manuscript, the replication package and journal preparation.
Boundaries of this service
A computer science and AI synthesis describes published empirical evidence. It does not build software or models, assess the security or safety of any system, or certify performance. Evidence ages quickly, many experiments use students, and benchmark results can be affected by tuning and data leakage, so findings should be read with those limits in mind.
Frequently asked questions
Can software engineering experiments be meta-analyzed?
Yes, when enough comparable experiments report means, standard deviations and sample sizes. Participants, tasks and tools are coded as moderators.
What is a systematic mapping study?
A review that classifies studies by topic, method and context to show where evidence exists, without estimating effects.
Can machine learning benchmark results be pooled?
Only with great care. Protocols, tuning and variance differ, so we usually tabulate and compare under matched conditions.
What is data leakage?
Information from the test set influencing training, which inflates performance. Reviews assess whether studies guarded against it.
Do findings about language models stay valid?
Often not for long. Results depend on model version and prompt, so they are dated and not generalized.
Do you build or certify software and models?
No. The service provides research and evidence-synthesis support only.
References
- Kitchenham BA, Charters S. Guidelines for performing systematic literature reviews in software engineering. EBSE Technical Report EBSE-2007-01. Keele University and University of Durham; 2007.
- Petersen K, Vakkalanka S, Kuzniarz L. Guidelines for conducting systematic mapping studies in software engineering: an update. Inf Softw Technol. 2015;64:1-18.
- Kapoor S, Narayanan A. Leakage and the reproducibility crisis in machine-learning-based science. Patterns. 2023;4(9):100804.
- Henderson P, Islam R, Bachman P, Pineau J, Precup D, Meger D. Deep reinforcement learning that matters. Proc AAAI Conf Artif Intell. 2018;32(1):3207-3214.
- Pineau J, Vincent-Lamarre P, Sinha K, et al. Improving reproducibility in machine learning research. J Mach Learn Res. 2021;22(164):1-20.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.