Meta-analysis and evidence synthesis for biomedical engineering

Biomedical engineering applies engineering to medicine and biology, from implants and imaging to wearable sensors and tissue engineering. Its evidence includes bench tests, animal studies, device trials and validation studies that compare a measurement against a reference. Reviews have to handle device iterations, operator learning curves, agreement statistics and the different levels of evidence behind regulatory approval.

Evidence synthesis in biomedical engineering

Biomedical engineering produces devices and methods used in diagnosis, treatment and monitoring: orthopedic and cardiac implants, imaging systems, prosthetics, wearable sensors, drug delivery systems and engineered tissues. Evidence about them ranges from bench tests to randomized trials. Systematic reviews address questions such as how accurate a wearable heart-rate monitor is, whether a new implant design has lower revision rates, how well an imaging algorithm detects disease, and how a biomaterial performs in animals before human use.

Several features set this evidence apart from drug trials. Devices change between versions, so evidence about an earlier model may not apply to the current one. Results depend on the skill of the operator, with learning curves that affect early outcomes. Blinding is often impossible. Many devices reach the market on the basis of equivalence to earlier devices, with limited clinical evidence. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for engineering and technology and connects with the medical and health sciences area. It gives no clinical advice.

Levels of evidence for devices

Types of device evidence
LevelTypical contentIssue for synthesis
Bench and simulationMechanical, electrical or computational testsStandardized but not clinical; conditions differ from use
Animal and in vitroBiocompatibility, function, healingSpecies differences; small samples; reporting quality
Validation studyMeasurement compared with a reference methodAgreement statistics, not clinical outcomes; spectrum of participants
Clinical trialRandomized or non-randomized comparisonBlinding often impossible; learning curve; rapid device change
Registry or post-marketLongitudinal data on use and failureObservational; reporting varies; useful for rare failures and long-term performance

A review should keep these levels apart, since each answers a different question. A device that performs well on the bench may not benefit patients. Registry data, such as joint replacement registries, contribute long-term revision rates for thousands of implants and are valuable for safety signals, though they are subject to confounding by indication and incomplete follow-up. Reviews should describe the evidence at each level and be explicit about the gaps.

Agreement and measurement validation

Many devices, such as blood pressure monitors, pulse oximeters and wearable sensors, are validated by comparing their readings with a reference method. The correlation between two methods is a poor measure of agreement, because methods can correlate strongly while differing systematically. The Bland-Altman approach plots the difference against the average and reports the mean difference (bias) and limits of agreement. Meta-analysis of agreement studies pools the bias and its variability across studies, using methods that handle both the mean difference and the standard deviation of differences, and that need information often missing from reports. A review should record the reference method, the range of values tested, the conditions of measurement (rest, exercise, skin tone, motion) and the population, since accuracy often differs across these.

Standards for acceptable error exist for some device types, and a review can report the proportion of studies meeting them. It should not state that a device is accurate in general without the conditions that were tested.

Learning curves, operators and iteration

Outcomes of surgical devices and procedures depend on operator experience. Early cases show more complications, and studies that include early adopters may show worse results than those of experienced users. Trials comparing a new device with the standard one can be biased if one is used by experts and the other by novices. A review should extract operator experience and case volume where reported, test them as moderators, and note whether the trial design handled learning. Because devices are revised, a review should record the model or generation, and avoid pooling results across designs with different performance without a moderator. The IDEAL framework describes stages of evaluation for surgical and device innovations and informs how to interpret early evidence.

Imaging and algorithms

Studies of imaging systems and algorithms report sensitivity, specificity and AUC compared with a reference standard. The methods on the diagnostic accuracy page apply, including bivariate models and risk-of-bias assessment with QUADAS-2. For algorithms, further concerns include training on data from one site, retrospective test sets, and lack of external validation. A review should record whether performance was tested on independent data from other centers and how many studies did so. Comparisons of algorithms with clinicians depend on the reading conditions and the case mix.

Biomaterials and tissue engineering

Biomaterials and tissue engineering studies are often preclinical, in animals, and report outcomes such as bone formation, tissue integration or inflammation. Reporting of randomization, blinding and sample size is often weak, effects are heterogeneous and publication bias is probable. The methods for preclinical meta-analysis apply, including risk-of-bias tools for animal studies, standardized mean differences and exploration of heterogeneity by species, model and material. Findings are suggestive of promise or problems and do not show efficacy in humans, and the review says so in its conclusions.

Regulatory evidence and post-market data

Regulators in different regions apply different rules to devices, based on risk class, and many lower-risk devices reach the market on the basis of similarity to earlier devices. Summaries of regulatory documents, where public, can show what evidence supported approval, but the documents differ by jurisdiction, and a review should not assume that approval implies effectiveness. Post-market surveillance, complaint databases and recalls give safety signals but have under-reporting and cannot estimate rates without denominators. This service does not prepare regulatory submissions or advise on compliance.

Publication bias and industry sponsorship

Many device trials are sponsored by manufacturers, and industry sponsorship has been associated with more favorable conclusions in some studies. We record funding and conflicts, compare sponsored and independent studies, check small-study effects and search registries and regulatory sources to identify unpublished studies.

Wearables, skin tone and real-world conditions

Consumer and clinical wearables estimate heart rate, oxygen saturation, sleep, steps and energy expenditure. Validation studies are common but vary in quality and often use small samples of healthy young adults in laboratory settings. Accuracy tends to fall during movement, in people with darker skin for optical sensors in some designs, and at extremes of the measured range. Reviews of wearables have found that step counts are usually accurate at walking speeds but less so at slow speeds, and that energy expenditure estimates have large errors. A review should extract the device model and firmware, the participant characteristics (including skin tone when reported), the activity condition and the reference method, and report accuracy by condition.

Software updates change device behavior without changing the hardware, so a study on one version may not describe the current product. Reviews state the date and version information they found and treat undated claims with caution, because a validation published several years ago may describe sensors and algorithms that have since been replaced.

Rehabilitation engineering, prosthetics and function

Studies of prosthetics, exoskeletons and assistive devices measure gait, energy cost, function and user satisfaction. Samples are small and heterogeneous in the level of amputation or impairment, and devices are fitted and adjusted for each user, so the intervention is the device plus the fitting process. Outcomes of performance tests in the laboratory may not predict use in daily life, and users' own priorities, such as comfort, appearance and reliability, may not be captured by performance measures. A review should record the population, the fitting and training provided, the outcome measures and whether users' views were sought. See also the rehabilitation page for clinical outcomes.

Common pitfalls we look for

  • Using correlation as a measure of agreement.
  • Mixing device versions without a moderator.
  • Ignoring operator experience and learning curves.
  • Treating bench performance as clinical benefit.
  • Accepting algorithm accuracy without external validation.
  • Assuming that regulatory approval shows effectiveness.

Planning a biomedical engineering synthesis

We help define the device or method, the clinical or measurement context, the comparator and outcomes, plan searches in MEDLINE, Embase, IEEE Xplore, Scopus, Web of Science and device registries, and set up coding of device version, operator experience, reference standard, conditions, sponsorship and level of evidence. See the systematic review service for scope and process.

An invented example of limits of agreement

Suppose a wrist sensor reads heart rate 1.0 beat per minute higher than a reference on average, with a standard deviation of differences of 4.0. The limits of agreement are the bias plus and minus 1.96 times 4.0, which gives about minus 6.8 to 8.8 beats per minute. A tool with a small average bias can still be wrong by almost 9 beats in a single reading, and whether this matters depends on the use. During exercise, the standard deviation might double to 8.0, widening the limits to about minus 14.7 to 16.7. A review of such sensors would report bias and limits under each condition and would not say that the sensor is accurate without saying when.

Coding and transparency

Coding frames record device name and version, manufacturer, technology type, reference standard, population and spectrum, conditions of use, operator experience, outcome and unit, level of evidence, sponsorship and country. Two coders extract data independently on a sample, and the coded data and scripts are shared with the final report.

How we support research projects in this area

Support

From a technical question to a published review

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    Help choosing between a mapping study, a systematic review and a meta-analysis, and a protocol with research questions.

  • Searching and extraction

    Database search, snowballing and grey literature, with screening and a classification scheme.

  • Synthesis

    Structured synthesis, quality assessment and, where experiments are comparable, meta-analysis.

  • Manuscript and submission

    The manuscript, the replication package and journal preparation.

Get a quoteDescribe the technology, the kinds of studies and the target journal or conference.

Boundaries of this service

A biomedical engineering synthesis describes published evidence on devices and methods. It does not design devices, prepare regulatory submissions, select devices or give clinical advice. Devices change between versions, results depend on operators and conditions, and bench or validation evidence does not show clinical benefit.

Frequently asked questions

Why is correlation not enough for validating a device?

Two methods can correlate strongly yet differ systematically. Agreement analysis reports bias and limits of agreement.

How do you handle learning curves?

By extracting operator experience and case volume, testing them as moderators and noting whether trials accounted for learning.

Can different device versions be pooled?

Only with a moderator for version. Redesigns can change performance, so evidence on one model may not apply to another.

Does regulatory approval show a device works?

Not necessarily. Many devices are approved by equivalence to earlier ones, with limited clinical evidence.

Are algorithm accuracy results reliable?

Only if tested on independent external data. Reviews report how many studies did so.

Do you design devices or advise on clinical use?

No. The service provides research and evidence-synthesis support only.

References

  1. Bland JM, Altman DG. Statistical methods for assessing agreement between two methods of clinical measurement. Lancet. 1986;1(8476):307-310.
  2. McCulloch P, Altman DG, Campbell WB, et al. No surgical innovation without evaluation: the IDEAL recommendations. Lancet. 2009;374(9695):1105-1112.
  3. Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
  4. Hooijmans CR, Rovers MM, de Vries RBM, Leenaars M, Ritskes-Hoitinga M, Langendam MW. SYRCLE's risk of bias tool for animal studies. BMC Med Res Methodol. 2014;14:43.
  5. Sedrakyan A, Campbell B, Merino JG, Kuntz R, Hirst A, McCulloch P. IDEAL-D: a rational framework for evaluating and regulating the use of medical devices. BMJ. 2016;353:i2372.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.