Evidence synthesis in surgery
Surgical practice has often changed on the basis of case series and expert opinion, with randomized trials arriving later and more slowly than in medicine. Reasons include difficulty in recruiting patients who must choose between operations, the lack of equipoise among surgeons, the rapid change of techniques and the cost of trials. As a result, systematic reviews in surgery usually combine a small number of randomized trials with many observational studies, and the question of whether and how to combine them is a central methodological decision.
The field also has strong traditions of quality improvement and audit. Registries such as national surgical databases contain information on very large numbers of patients and can answer questions about rare complications and about outcomes in routine care. They are open to confounding and variable data quality. The IDEAL framework, proposed by McCulloch and colleagues in 2009, describes stages in the evaluation of surgical innovation, from the first use in humans through development, exploration, assessment and long-term study. It helps to decide what level of evidence can be expected at each stage.
The general framework is in meta-analysis and systematic review, adapted to surgery as follows.
Randomized and observational evidence
| Source | Strength | Main limit |
|---|---|---|
| Randomized trials | Protect against confounding when well conducted | Small, few, surgeon-dependent, blinding limited, selective centers |
| Sham-controlled trials | Test the specific effect of a procedure | Ethical limits; few exist; sham may not be credible |
| Cohort and registry studies | Large samples, real-world practice, rare outcomes | Confounding by indication, variable definitions, missing data |
| Case series | Describe feasibility and complications of new procedures | No comparison group; selected patients; publication bias |
A review can analyze randomized and observational evidence separately and compare the results, or restrict the main analysis to randomized trials and use observational data for harms. Pooling the two in one estimate hides the difference in bias. Confounding by indication is the central concern: surgeons choose the operation for patients according to risk, so patients having one approach differ from those having another. Adjusted estimates and propensity-score-based analyses reduce but do not eliminate this. Risk of bias is assessed with RoB 2 for trials and ROBINS-I for non-randomized comparisons.
Blinding, surgeon effects and the learning curve
Neither patients nor surgeons can usually be blinded to an operation, so subjective outcomes such as pain and satisfaction may be influenced by expectations. Blinding of outcome assessors and patients is possible in some trials, for example when wounds are covered or when anesthesia is used to conceal the approach, and reviews should record whether it was done. The effect of unblinding is largest for patient-reported outcomes and smallest for hard outcomes such as death.
Surgeon skill and experience influence results. A new technique performed early on the learning curve shows more complications than the same technique performed by experienced teams. Trials that require surgeons to have completed a number of procedures reduce this effect but also limit generalizability, because results from expert centers may not apply in smaller hospitals. A review should record the experience required, the number of surgeons and centers, and whether the analysis accounts for clustering by surgeon. Meta-regression on trial year or center volume can examine whether results depend on experience. Technique changes over time also mean that older trials may not describe current practice.
Complications, harms and proportions
Complications are central to surgical evidence. Classification systems such as Clavien-Dindo grade complications by the intervention needed to treat them, which allows reviews to separate minor from major events. Reviews should extract the grade of complication and avoid pooling "any complication" with "major complication" without distinction. Event rates are often low, and some studies report zero events. The guide on zero-event studies explains methods for sparse data.
When no comparison group exists, a review may pool the proportion of patients with an outcome in single-arm series. As an illustration, a series with 400 patients and 36 complications has a proportion of 0.09. The Wald 95 percent confidence interval is 0.090 plus or minus 1.96 x 0.0143, which gives 0.062 to 0.118. The numbers are invented. With small series or proportions near 0 or 1, a logit or Freeman-Tukey transformation, or a generalized linear mixed model, gives better behavior than the raw proportion, as covered in prevalence meta-analysis. Pooled single-arm proportions describe outcomes in treated patients, not the effect of surgery, and heterogeneity is usually large because of patient selection.
Patient-reported outcomes and function
Surgery for non-fatal conditions, such as joint replacement, hernia repair or bariatric surgery, is judged by function, pain, quality of life and satisfaction. Instruments vary, and results are expressed as standardized mean differences or changes in scale points. Reviews should state the minimal important difference where available and avoid claiming benefit when the difference is smaller. Follow-up is often too short to show durable results or late failure, such as implant loosening or hernia recurrence. Recurrence and reoperation need time-to-event methods or at least a stated follow-up period.
Comparisons of surgery with non-operative care, such as physiotherapy for knee pain or watchful waiting for hernia, are a growing area. In these trials crossover is frequent, and intention-to-treat estimates dilute the surgical effect, while per-protocol estimates are biased. Reviews should extract crossover rates and report both analyses when available.
New techniques, devices and innovation
New techniques, including laparoscopic, robotic and endovascular approaches, are often adopted before trials are done. Early evidence consists of case series from enthusiastic centers, with selected patients and favorable reporting. Later randomized trials can show smaller benefits, equivalence or harm. A review of an innovation should say where it stands on the IDEAL framework and should treat early series as hypothesis generating. Industry involvement in device trials is common, and the funding and the role of manufacturers in design and analysis should be recorded. Devices change in generations, and pooled results from different generations of a device may not describe the current product.
Perioperative care and enhanced recovery
A large share of surgical research concerns care around the operation: anesthesia techniques, antibiotic prophylaxis, wound closure, drains, nutrition support, analgesia and enhanced recovery programs that bundle many measures. Trials of these measures are usually easier to conduct and to blind than trials of operations, and they provide much of the randomized evidence in surgical journals. Bundled programs raise the question of which element matters. Component analysis, as in component network meta-analysis, can help when trials vary the components independently, although the number of trials is often too small.
Outcomes include surgical site infection, length of stay, return of bowel function and complications. Length of stay is skewed and depends on local discharge policy, so means can be unstable, and reviews should consider the median or a ratio of means and note that policy differences between countries and eras limit comparability. Infection definitions differ between trials, and surveillance after discharge affects the rate recorded.
Interpreting results for hospitals and patients
Surgeons and health services need to know how large a difference is in absolute terms, how it varies by patient risk and how it depends on the setting. A relative reduction in complications of 20 percent means a lot for a procedure with a 20 percent complication rate and little for one with a 1 percent rate. Reviews should present absolute effects at stated baseline risks, as with other health topics. Patients also weigh recovery time, pain, appearance of scars and return to work, which are less often measured. Where such outcomes are missing, the review should say so, since a conclusion about the best operation based only on complications may miss what matters to patients.
Applicability is a concern. Trials often exclude older patients, those with several illnesses and those needing emergency surgery, yet these groups make up a large share of operations in practice. Registry data can supplement trials for these patients, with caution on confounding. A review should list the exclusion criteria of the trials and judge how far the results apply to patients who would be excluded.
Common pitfalls we look for
- Pooling randomized and observational studies in one estimate without comparing them.
- Ignoring surgeon and center clustering and learning-curve effects.
- Combining minor and major complications in a single outcome.
- Treating single-arm proportions as treatment effects.
- Reporting only intention-to-treat or only per-protocol results when crossover is high.
- Applying results from expert centers to all hospitals without comment.
Planning and reporting
The protocol defines the procedures and comparators, the patient population, the outcomes with their definitions and time points, the designs to be included and how they will be analyzed, and the planned analyses for surgeon experience and era. Searches cover MEDLINE, Embase, CENTRAL and trial registries, with conference abstracts since many surgical studies appear there first. Two reviewers screen and extract. Reporting follows PRISMA 2020, and MOOSE for meta-analyses of observational studies, with a protocol registered in PROSPERO. Certainty is rated with GRADE, with observational studies starting at low certainty unless upgraded.
How we support research projects in this area
From a clinical question to a published review
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A structured question, eligibility criteria and an analysis plan, with registration prepared where appropriate.
Searching and extraction
Search strategies for the relevant databases and registries, screening and data extraction, and risk-of-bias assessment by design.
Synthesis
Pairwise, network, diagnostic accuracy, prognostic or dose-response analysis, with a GRADE assessment for each outcome.
Manuscript and submission
Reporting-guideline checklists, the manuscript and the preparation of submission materials.
Boundaries of this service
A review of surgical studies describes outcomes in groups of patients. It does not tell a person whether to have an operation, which depends on their condition, other illness, preferences and the advice of their surgeon. We do not provide surgical recommendations, assessment of operative risk for an individual or advice on the care of any patient.
Frequently asked questions
Can I combine randomized trials and observational studies?
They can be shown together, but pooling them in one estimate hides different biases. Many reviews analyze them separately and compare the results.
How do I deal with surgeon experience?
Record the experience required and the number of surgeons, account for clustering where possible and use meta-regression on era or volume.
How should complications be analyzed?
By grade, using a classification such as Clavien-Dindo, with methods for sparse data when events are rare.
Can I pool single-arm series?
Yes, as proportions, using suitable transformations or mixed models. They describe outcomes in treated patients and not the effect of surgery.
What is the IDEAL framework?
A staged framework for evaluating surgical innovation, from first-in-human use to long-term study.
Do you give surgical advice?
No. The service provides research and evidence-synthesis support only.
References
- McCulloch P, Altman DG, Campbell WB, et al. No surgical innovation without evaluation: the IDEAL recommendations. Lancet. 2009;374(9695):1105-1112.
- Sterne JA, Hernan MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
- Dindo D, Demartines N, Clavien PA. Classification of surgical complications: a new proposal with evaluation in a cohort of 6336 patients and results of a survey. Ann Surg. 2004;240(2):205-213.
- Stroup DF, Berlin JA, Morton SC, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-2012.
- Schwarzer G, Chemaitelly H, Abu-Raddad LJ, Rucker G. Seriously misleading results using inverse of Freeman-Tukey double arcsine transformation in meta-analysis of single proportions. Res Synth Methods. 2019;10(3):476-483.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.