Evidence synthesis in biomedical engineering
Biomedical engineering produces devices and methods used in diagnosis, treatment and monitoring: orthopedic and cardiac implants, imaging systems, prosthetics, wearable sensors, drug delivery systems and engineered tissues. Evidence about them ranges from bench tests to randomized trials. Systematic reviews address questions such as how accurate a wearable heart-rate monitor is, whether a new implant design has lower revision rates, how well an imaging algorithm detects disease, and how a biomaterial performs in animals before human use.
Several features set this evidence apart from drug trials. Devices change between versions, so evidence about an earlier model may not apply to the current one. Results depend on the skill of the operator, with learning curves that affect early outcomes. Blinding is often impossible. Many devices reach the market on the basis of equivalence to earlier devices, with limited clinical evidence. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for engineering and technology and connects with the medical and health sciences area. It gives no clinical advice.
Levels of evidence for devices
| Level | Typical content | Issue for synthesis |
|---|---|---|
| Bench and simulation | Mechanical, electrical or computational tests | Standardized but not clinical; conditions differ from use |
| Animal and in vitro | Biocompatibility, function, healing | Species differences; small samples; reporting quality |
| Validation study | Measurement compared with a reference method | Agreement statistics, not clinical outcomes; spectrum of participants |
| Clinical trial | Randomized or non-randomized comparison | Blinding often impossible; learning curve; rapid device change |
| Registry or post-market | Longitudinal data on use and failure | Observational; reporting varies; useful for rare failures and long-term performance |
A review should keep these levels apart, since each answers a different question. A device that performs well on the bench may not benefit patients. Registry data, such as joint replacement registries, contribute long-term revision rates for thousands of implants and are valuable for safety signals, though they are subject to confounding by indication and incomplete follow-up. Reviews should describe the evidence at each level and be explicit about the gaps.
Agreement and measurement validation
Many devices, such as blood pressure monitors, pulse oximeters and wearable sensors, are validated by comparing their readings with a reference method. The correlation between two methods is a poor measure of agreement, because methods can correlate strongly while differing systematically. The Bland-Altman approach plots the difference against the average and reports the mean difference (bias) and limits of agreement. Meta-analysis of agreement studies pools the bias and its variability across studies, using methods that handle both the mean difference and the standard deviation of differences, and that need information often missing from reports. A review should record the reference method, the range of values tested, the conditions of measurement (rest, exercise, skin tone, motion) and the population, since accuracy often differs across these.
Standards for acceptable error exist for some device types, and a review can report the proportion of studies meeting them. It should not state that a device is accurate in general without the conditions that were tested.
Learning curves, operators and iteration
Outcomes of surgical devices and procedures depend on operator experience. Early cases show more complications, and studies that include early adopters may show worse results than those of experienced users. Trials comparing a new device with the standard one can be biased if one is used by experts and the other by novices. A review should extract operator experience and case volume where reported, test them as moderators, and note whether the trial design handled learning. Because devices are revised, a review should record the model or generation, and avoid pooling results across designs with different performance without a moderator. The IDEAL framework describes stages of evaluation for surgical and device innovations and informs how to interpret early evidence.
Imaging and algorithms
Studies of imaging systems and algorithms report sensitivity, specificity and AUC compared with a reference standard. The methods on the diagnostic accuracy page apply, including bivariate models and risk-of-bias assessment with QUADAS-2. For algorithms, further concerns include training on data from one site, retrospective test sets, and lack of external validation. A review should record whether performance was tested on independent data from other centers and how many studies did so. Comparisons of algorithms with clinicians depend on the reading conditions and the case mix.
Biomaterials and tissue engineering
Biomaterials and tissue engineering studies are often preclinical, in animals, and report outcomes such as bone formation, tissue integration or inflammation. Reporting of randomization, blinding and sample size is often weak, effects are heterogeneous and publication bias is probable. The methods for preclinical meta-analysis apply, including risk-of-bias tools for animal studies, standardized mean differences and exploration of heterogeneity by species, model and material. Findings are suggestive of promise or problems and do not show efficacy in humans, and the review says so in its conclusions.
Regulatory evidence and post-market data
Regulators in different regions apply different rules to devices, based on risk class, and many lower-risk devices reach the market on the basis of similarity to earlier devices. Summaries of regulatory documents, where public, can show what evidence supported approval, but the documents differ by jurisdiction, and a review should not assume that approval implies effectiveness. Post-market surveillance, complaint databases and recalls give safety signals but have under-reporting and cannot estimate rates without denominators. This service does not prepare regulatory submissions or advise on compliance.
Publication bias and industry sponsorship
Many device trials are sponsored by manufacturers, and industry sponsorship has been associated with more favorable conclusions in some studies. We record funding and conflicts, compare sponsored and independent studies, check small-study effects and search registries and regulatory sources to identify unpublished studies.
Wearables, skin tone and real-world conditions
Consumer and clinical wearables estimate heart rate, oxygen saturation, sleep, steps and energy expenditure. Validation studies are common but vary in quality and often use small samples of healthy young adults in laboratory settings. Accuracy tends to fall during movement, in people with darker skin for optical sensors in some designs, and at extremes of the measured range. Reviews of wearables have found that step counts are usually accurate at walking speeds but less so at slow speeds, and that energy expenditure estimates have large errors. A review should extract the device model and firmware, the participant characteristics (including skin tone when reported), the activity condition and the reference method, and report accuracy by condition.
Software updates change device behavior without changing the hardware, so a study on one version may not describe the current product. Reviews state the date and version information they found and treat undated claims with caution, because a validation published several years ago may describe sensors and algorithms that have since been replaced.
Rehabilitation engineering, prosthetics and function
Studies of prosthetics, exoskeletons and assistive devices measure gait, energy cost, function and user satisfaction. Samples are small and heterogeneous in the level of amputation or impairment, and devices are fitted and adjusted for each user, so the intervention is the device plus the fitting process. Outcomes of performance tests in the laboratory may not predict use in daily life, and users' own priorities, such as comfort, appearance and reliability, may not be captured by performance measures. A review should record the population, the fitting and training provided, the outcome measures and whether users' views were sought. See also the rehabilitation page for clinical outcomes.
Common pitfalls we look for
- Using correlation as a measure of agreement.
- Mixing device versions without a moderator.
- Ignoring operator experience and learning curves.
- Treating bench performance as clinical benefit.
- Accepting algorithm accuracy without external validation.
- Assuming that regulatory approval shows effectiveness.
Planning a biomedical engineering synthesis
We help define the device or method, the clinical or measurement context, the comparator and outcomes, plan searches in MEDLINE, Embase, IEEE Xplore, Scopus, Web of Science and device registries, and set up coding of device version, operator experience, reference standard, conditions, sponsorship and level of evidence. See the systematic review service for scope and process.
An invented example of limits of agreement
Suppose a wrist sensor reads heart rate 1.0 beat per minute higher than a reference on average, with a standard deviation of differences of 4.0. The limits of agreement are the bias plus and minus 1.96 times 4.0, which gives about minus 6.8 to 8.8 beats per minute. A tool with a small average bias can still be wrong by almost 9 beats in a single reading, and whether this matters depends on the use. During exercise, the standard deviation might double to 8.0, widening the limits to about minus 14.7 to 16.7. A review of such sensors would report bias and limits under each condition and would not say that the sensor is accurate without saying when.
Coding and transparency
Coding frames record device name and version, manufacturer, technology type, reference standard, population and spectrum, conditions of use, operator experience, outcome and unit, level of evidence, sponsorship and country. Two coders extract data independently on a sample, and the coded data and scripts are shared with the final report.
How we support research projects in this area
From a technical question to a published review
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
Help choosing between a mapping study, a systematic review and a meta-analysis, and a protocol with research questions.
Searching and extraction
Database search, snowballing and grey literature, with screening and a classification scheme.
Synthesis
Structured synthesis, quality assessment and, where experiments are comparable, meta-analysis.
Manuscript and submission
The manuscript, the replication package and journal preparation.
Boundaries of this service
A biomedical engineering synthesis describes published evidence on devices and methods. It does not design devices, prepare regulatory submissions, select devices or give clinical advice. Devices change between versions, results depend on operators and conditions, and bench or validation evidence does not show clinical benefit.
Frequently asked questions
Why is correlation not enough for validating a device?
Two methods can correlate strongly yet differ systematically. Agreement analysis reports bias and limits of agreement.
How do you handle learning curves?
By extracting operator experience and case volume, testing them as moderators and noting whether trials accounted for learning.
Can different device versions be pooled?
Only with a moderator for version. Redesigns can change performance, so evidence on one model may not apply to another.
Does regulatory approval show a device works?
Not necessarily. Many devices are approved by equivalence to earlier ones, with limited clinical evidence.
Are algorithm accuracy results reliable?
Only if tested on independent external data. Reviews report how many studies did so.
Do you design devices or advise on clinical use?
No. The service provides research and evidence-synthesis support only.
References
- Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet. 1986;1(8476):307-310.
- McCulloch P, Altman DG, Campbell WB, et al. No surgical innovation without evaluation: the IDEAL recommendations. Lancet. 2009;374(9695):1105-1112.
- Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
- Hooijmans CR, Rovers MM, de Vries RBM, Leenaars M, Ritskes-Hoitinga M, Langendam MW. SYRCLE's risk of bias tool for animal studies. BMC Med Res Methodol. 2014;14:43.
- Sedrakyan A, Campbell B, Merino JG, Kuntz R, Hirst A, McCulloch P. IDEAL-D: a rational framework for evaluating and regulating the use of medical devices. BMJ. 2016;353:i2372.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.