Evidence synthesis in engineering
Software engineering adopted systematic review early. Kitchenham and Charters published guidelines for systematic literature reviews in software engineering in 2007, adapted from medical methods, and the field now publishes hundreds of systematic reviews and mapping studies each year. Other engineering fields use similar methods, from civil and energy engineering to human-computer interaction, biomedical engineering and artificial intelligence.
The evidence base differs from that in health research. Many papers propose a new technique or tool and evaluate it on a benchmark chosen by the authors. Comparisons between papers are hard because datasets, hardware, parameters and metrics differ. Controlled experiments with human participants, such as comparisons of programming languages or design methods, exist but are fewer, small and often done with students. Industrial case studies are rich in context but not generalizable by statistics. And technology changes quickly, so evidence ages fast.
A synthesis should therefore begin with the question and the type of evidence. Not every question needs a meta-analysis, and some are better served by a map of the field. Our guidance on review types applies here.
Choosing the form of synthesis
| Form | Question it answers | Typical output |
|---|---|---|
| Systematic mapping study | What research exists on a topic, and how is it distributed? | Classification scheme, counts, trends, gaps (Petersen et al. 2015) |
| Systematic literature review | What does the evidence say about a specific question? | Structured comparison of techniques, with quality assessment |
| Meta-analysis of experiments | What is the average effect of a technique, and does it vary? | Pooled effect size with heterogeneity, when experiments are comparable |
| Tertiary study | What do existing reviews show about a field? | Review of reviews, quality of reviews and gaps |
| Grey literature review | What do practitioners report? | Findings from blogs, white papers and reports, appraised for credibility |
Mapping studies are very common in engineering. They classify papers by technique, application, evaluation method and year, and show where research is concentrated. They do not judge which technique is best. For that purpose a systematic literature review with explicit quality criteria is needed, and a meta-analysis only if the studies measure the same outcome in a comparable way.
Searching: databases, snowballing and grey literature
Engineering literature is spread across IEEE Xplore, ACM Digital Library, Scopus, Web of Science, Engineering Village (Compendex), SpringerLink and ScienceDirect, plus arXiv and conference proceedings that are not always well indexed. Search strings have to deal with the inconsistent terminology of technology fields, and pilot searches against a set of known papers help to test the string. Wohlin proposed guidelines for snowballing, in which the reference lists and citing papers of a starting set are followed iteratively, and many engineering reviews use snowballing alongside or instead of database search.
Grey literature matters because practice often leads research: white papers, standards, technical reports and blog posts by practitioners contain evidence not found in journals. Including it requires criteria for credibility, such as the authority of the source, the method described and the transparency of data. The protocol should state these criteria and how grey sources will be weighed.
Quality assessment and bias in technology studies
Quality criteria in engineering reviews look at whether a study states its research questions, describes the context, defines its metrics, compares with a suitable baseline and discusses threats to validity. Kitchenham and Charters give examples. For controlled experiments, standard concerns apply: participant selection, random allocation, task realism and the risk that the authors who developed a technique also evaluate it. Studies by the developers of a technique tend to report better results than independent evaluations, so the relationship of authors to the technique is worth coding.
Benchmark studies have a particular risk of overfitting to a standard dataset and of leakage between training and test data. Machine-learning reviews should note whether studies used independent test sets, whether hyperparameters were tuned on test data, whether results were repeated across random seeds and whether code and data are available. Without these, reported differences between methods may not hold up.
When meta-analysis is appropriate
Meta-analysis of experiments has been used in software engineering for topics such as test-driven development, pair programming and requirements elicitation. Miller discussed the application of meta-analytic procedures to software experiments in 2000. The conditions are the same as elsewhere: studies must measure comparable outcomes under comparable conditions and report enough data to compute an effect size and variance. These conditions are often unmet. Experiments use different tasks, participants and measures of quality or productivity, and heterogeneity is large.
When pooling is feasible, standard methods apply: standardized mean differences, random-effects models, meta-regression on features such as participant experience, and tests for small-study effects. When it is not, the review should present a structured comparison, such as a table of results by technique and context, and say clearly why pooling was not done. Reporting that the evidence is too heterogeneous to combine is a result.
Benchmarks pose a particular problem. Comparing reported scores from different papers on the same dataset appears simple, but papers differ in splits, preprocessing and tuning. A fair comparison needs reimplementation under the same conditions, which is a primary study and not a review.
Reporting, replicability and ageing evidence
Engineering reviews should publish the search strings, the inclusion and exclusion criteria, the list of included papers, the data extraction sheet and the classification scheme, ideally in an open repository. This allows replication and later updating. Because technology changes quickly, the date of the last search and the period covered should be stated, and a review should say how long its conclusions are likely to remain useful. For fast-moving topics, a living map or an updatable database is a better format than a static paper.
Reporting can follow PRISMA 2020 where the review is similar to a systematic review in the health sense, and the specific guidelines of the engineering field for mapping studies and quality assessment. The result should present the findings in terms that practitioners and researchers can use, with a clear statement of the evidence behind each.
Specialties and sub-fields
Sub-fields differ in the kinds of studies and the feasibility of pooling. Pages for sub-fields are added as they are completed.
A worked example of a mapping result
Suppose a mapping study of a testing technique screens its search results and includes 120 papers. The authors classify them by evaluation method. The counts are invented to show how a mapping result looks.
| Evaluation method | Papers | Share |
|---|---|---|
| Controlled experiment | 18 | 15.0% |
| Benchmark comparison | 54 | 45.0% |
| Industrial case study | 22 | 18.3% |
| Simulation | 16 | 13.3% |
| Survey or interview | 10 | 8.3% |
The map shows that nearly half of the papers evaluate the technique on a benchmark and only 15 percent use a controlled experiment. A reader can conclude that evidence of the kind needed for a meta-analysis is scarce, and that a pooled effect from the 18 experiments would rest on a small and perhaps unrepresentative part of the field. A mapping study stops here: it reports the distribution and the gaps. A follow-up review might examine the 18 experiments in detail, assess their quality and decide whether they are comparable enough to pool. The map also helps to plan new research, since it shows that few studies are done in industrial settings with practitioners.
Experiments with people
Many software and design experiments use students as participants because they are easy to recruit. Whether the results apply to professionals is debated, and a review should record the participants' experience and test whether it moderates the effect. Tasks are often small and artificial. Measures such as productivity, defect counts or time to complete a task depend on the task and the tools, so standardizing effects across studies is only meaningful when tasks are similar. Where several experiments are replications by one group, they are not independent, and the review should model the group as a source of dependence. Preregistration and registered reports are becoming more common in empirical software engineering, and reviews can compare registered with unregistered studies.
A typical workflow and how long it takes
A systematic literature review in engineering follows a sequence: a protocol with research questions and criteria, a pilot search, the full search across databases and snowballing, de-duplication, screening of titles, abstracts and full texts by two reviewers, extraction into a standard form, quality assessment, synthesis and reporting. Mapping studies follow the same sequence without the quality assessment and with a classification scheme developed from the papers. For a topic with a few hundred included papers, the work usually takes several months of part-time effort by a small team. A tertiary study that reviews existing reviews is shorter.
Because the field moves quickly, the protocol should set a search date and plan for updates. A review that took two years to complete may describe a field that has since changed. For fast-moving topics a rapid review can be considered, with limits that are stated. The review team should also include at least one member with practical experience in the field, who can judge whether the classification scheme and the research questions match how practitioners see the problem and the way terms are used in industry today, which often differs from academic usage.
How we support research projects in this area
From a technical question to a published review
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
Help choosing between a mapping study, a systematic review and a meta-analysis, and a protocol with research questions.
Searching and extraction
Database search, snowballing and grey literature, with screening and a classification scheme.
Synthesis
Structured synthesis, quality assessment and, where experiments are comparable, meta-analysis.
Manuscript and submission
The manuscript, the replication package and journal preparation.
Boundaries of this service
A synthesis of engineering studies describes what the literature reports. It is not a design review or a safety certification, and it does not replace testing in the conditions where a technology will be used. Reported performance in a benchmark may not hold in practice. We do not provide engineering design, certification or professional engineering opinions.
Where studies cannot be pooled, we recommend a mapping study or a structured review and say so at the start.
Frequently asked questions
Is meta-analysis possible for engineering studies?
Sometimes, mainly for controlled experiments with comparable outcomes. Often the studies are too different, and a mapping study or structured review is more suitable.
What is a systematic mapping study?
A review that classifies the studies on a topic by technique, application and method to show how research is distributed, without judging which approach is best.
What is snowballing?
Following the reference lists and citations of a starting set of papers, repeatedly, to find further studies.
Should I include grey literature?
Often yes, because practice may lead research. Set credibility criteria in the protocol.
How do I compare machine-learning results from different papers?
Only after checking whether datasets, splits, preprocessing and tuning are comparable. Otherwise report the results side by side without pooling.
Do you provide engineering design or certification advice?
No. The service covers research and evidence-synthesis support only.
References
- Kitchenham B, Charters S. Guidelines for performing systematic literature reviews in software engineering. EBSE Technical Report EBSE-2007-01. Keele University and Durham University; 2007.
- Petersen K, Vakkalanka S, Kuzniarz L. Guidelines for conducting systematic mapping studies in software engineering: an update. Inf Softw Technol. 2015;64:1-18.
- Wohlin C. Guidelines for snowballing in systematic literature studies and a replication in software engineering. In: Proceedings of the 18th International Conference on Evaluation and Assessment in Software Engineering (EASE '14). ACM; 2014. Article 38.
- Miller J. Applying meta-analytical procedures to software engineering experiments. J Syst Softw. 2000;54(1):29-39.
- Cruzes DS, Dyba T. Research synthesis in software engineering: a tertiary study. Inf Softw Technol. 2011;53(5):440-455.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.