Meta-analysis in obstetrics and gynecology

Obstetric evidence concerns two patients at once, the mother and the baby, and often many outcomes at once. Gynecology covers surgery, fertility, cancer screening and menopause. Reviews in this field must handle pregnancy-specific outcomes, ethical limits on trials in pregnancy and heavy reliance on observational data for exposures.

Evidence synthesis in obstetrics and gynecology

Obstetrics has a distinguished place in the history of research synthesis. The first large Cochrane project, begun in the 1980s, was a database of controlled trials in pregnancy and childbirth, and reviews of corticosteroids for fetal lung maturity, magnesium sulfate for eclampsia and antenatal care showed how pooling trials could reveal effects that single trials missed. The Cochrane Pregnancy and Childbirth group remains one of the largest review groups.

The field has unusual features. Pregnant women are often excluded from drug trials, so evidence on drug safety in pregnancy comes mainly from observational studies. Outcomes are shared between mother and baby, and an intervention that benefits one may harm the other. Many outcomes are rare events, such as stillbirth or maternal death, so large samples or pooled data are needed. Outcomes are also inconsistently defined and reported, which led to the development of core outcome sets for several conditions. Gynecologic studies add questions about surgery, fertility treatment, screening and menopause.

The methods follow meta-analysis and systematic review, with the attention to pregnancy described here.

Pregnancy outcomes and their definitions

Common pregnancy outcomes and points to check
OutcomeTypical measurePoints to check
Preterm birthRisk ratioGestational age cutoff (37, 34, 32 weeks); spontaneous versus indicated
Stillbirth and neonatal deathRisk ratio, odds ratioRare events; definition by gestation or weight; competing risks
PreeclampsiaRisk ratioDiagnostic criteria changed over time; early versus late onset
Mode of birthRisk ratioCesarean rate depends on local practice and indication
Live birth (fertility)Risk ratio per woman or per cycleUnit of analysis: per woman, per cycle or per embryo transfer; cumulative results
Birth weightMean differenceAdjustment for gestational age; small for gestational age definitions

Definitions change across trials, countries and years, and a pooled estimate based on mixed definitions can mislead. Reviews should tabulate definitions and, where there are enough studies, analyze groups separately. Core outcome sets, developed through international consensus for conditions such as preeclampsia, preterm birth and infertility, have been created to reduce variation, and a review can compare the outcomes reported in trials with the core set to show gaps and selective reporting.

Unit of analysis: women, pregnancies, babies, cycles

Several levels of analysis exist in obstetrics and fertility research. A woman may have more than one pregnancy in a study, a pregnancy may have more than one baby (twins), and a woman in fertility treatment may have several cycles. Counting twins as independent babies, or cycles as independent women, overstates the sample size and narrows the interval. Reviews should record the unit used in each study and convert or adjust where possible. For fertility treatment, the live birth rate per woman randomized is the preferred outcome; rates per cycle or per transfer are less useful, because they depend on how many cycles were done and on the selection of embryos.

Multiple-birth pregnancies are a risk in themselves, so outcomes such as multiple pregnancy rate are essential in reviews of fertility treatments. Studies of interventions that raise the success rate by increasing multiple births do not show benefit overall.

Exposures in pregnancy and observational data

Evidence on medicines, infections, nutrition, environmental exposures and lifestyle in pregnancy comes largely from cohorts, registries and case-control studies. These studies are subject to confounding, including confounding by indication (women taking a drug have the condition it treats), recall bias in case-control designs and selection bias, for example when only live births are studied. Sibling comparisons and negative-control analyses help to address confounding, and reviews should note which studies used them. Pooling follows the MOOSE guidelines, with adjusted estimates, an assessment of risk of bias with ROBINS-E or similar tools, and a test of whether results differ by the method of adjustment.

The timing of exposure matters: effects in the first trimester differ from those in the third. Reviews should extract the time window and consider separate analyses. Absolute risks are important because the baseline risk of outcomes like congenital malformation is low, and a relative increase can still mean a small absolute risk. Presenting both helps women and clinicians to put findings in perspective.

Gynecologic surgery, screening and menopause

Gynecology includes surgery for fibroids, endometriosis and prolapse, cancer screening, contraception and menopausal symptoms. Surgical trials share the issues described for surgery: limited blinding, surgeon experience and variation in technique. Trials of menopausal symptoms use symptom diaries, with strong placebo responses. Screening studies, such as for cervical cancer, require attention to lead time and overdiagnosis, and to differences between the test accuracy and the effect on mortality. Contraceptive effectiveness is estimated with Pearl indices or life tables, and these measures differ in how they handle discontinuation, so they should not be pooled as if they were the same.

A worked reading of an absolute effect

Suppose a review finds that an intervention reduces preterm birth before 37 weeks, with a pooled risk ratio of 0.75. In the control groups 8% of women gave birth preterm. The numbers are invented to show the reading. In the intervention groups the expected rate is 0.08 x 0.75 = 0.06, or 6%, so the absolute reduction is 2 percentage points and the number needed to treat is 1/0.02 = 50.

A reader should then ask whether the women in the trials resemble the women in the question. If the trials enrolled women with a history of preterm birth, with a baseline risk of 8 percent or higher, the result does not apply to women at low risk, whose baseline may be 3 percent and whose absolute benefit would be about 0.75 percentage points. The reader should also ask what happened to the babies: reducing preterm birth matters because it is expected to reduce neonatal death and illness, and a review that reports the intervention's effects on those outcomes provides stronger evidence than one that reports only the intermediate outcome. Finally, harms to the mother and the baby should be shown beside benefits.

Prenatal screening, diagnosis and prediction

Prenatal screening for chromosomal conditions, gestational diabetes, preeclampsia and fetal growth restriction generates many diagnostic and prediction studies. Accuracy reviews use bivariate models for sensitivity and specificity, see diagnostic accuracy meta-analysis, and must account for the fact that the same test performs differently in populations with different baseline risk and at different gestational ages. For a screening test the useful quantities are the detection rate at a fixed false-positive rate and the predictive values in the population to be screened. Multivariable prediction models, such as models for preeclampsia, require external validation, and reviews should evaluate calibration as well as discrimination, using PROBAST to judge bias. Many published models are developed on small samples and are never validated in other settings.

Ultrasound measurements and biomarkers vary between laboratories and operators. Reviews should extract the thresholds used, since differences in cutoffs explain much of the variation between studies, and a summary receiver operating characteristic approach may be more informative than a single pooled pair.

Global health, equity and access to care

Most maternal and neonatal deaths occur in low- and middle-income countries, where the causes, the baseline risks and the available care differ from those in the settings of most trials. Interventions such as antenatal corticosteroids have been shown to work in hospitals with neonatal intensive care, and a large trial in low-resource settings found different results, which illustrates why the setting matters. Reviews should record where the trials were done, the level of care available and the baseline risks, and should avoid applying a pooled estimate from one setting to another without comment.

Disparities by ethnicity, income and geography are well documented in maternity care. The PRISMA-Equity extension encourages reviewers to consider such differences in the question, the analysis and the discussion. Where data allow, subgroup analyses by these factors can show whether benefits are shared. Participants in trials often differ from the people who bear the highest risk, and the gap should be reported.

Common pitfalls we look for

  • Counting twins or cycles as independent units.
  • Mixing definitions of preterm birth, preeclampsia or stillbirth.
  • Reporting an intermediate outcome without outcomes for the baby.
  • Applying results from high-risk women to low-risk women.
  • Treating observational associations as causal without attention to confounding by indication.
  • Ignoring harms to the mother or the baby.

Planning and reporting

The protocol defines the population (for example, singleton pregnancies, or women undergoing fertility treatment), the intervention and comparator, the outcomes with definitions for mother and baby, the unit of analysis and the planned subgroups by risk, parity and gestational age. Searches cover MEDLINE, Embase, CENTRAL, CINAHL, registries and, for global questions, regional databases. Risk of bias is assessed with RoB 2 and ROBINS-I. Certainty is rated with GRADE. Reporting follows PRISMA 2020, MOOSE for observational reviews and individual participant data extensions where used, with a protocol registered in PROSPERO.

How we support research projects in this area

Support

From a clinical question to a published review

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A structured question, eligibility criteria and an analysis plan, with registration prepared where appropriate.

  • Searching and extraction

    Search strategies for the relevant databases and registries, screening and data extraction, and risk-of-bias assessment by design.

  • Synthesis

    Pairwise, network, diagnostic accuracy, prognostic or dose-response analysis, with a GRADE assessment for each outcome.

  • Manuscript and submission

    Reporting-guideline checklists, the manuscript and the preparation of submission materials.

Get a quoteDescribe your question, study types and target journal.

Boundaries of this service

A review of obstetric and gynecologic studies describes outcomes in groups of women and babies. It does not tell a pregnant woman or a patient what to do, which depends on her health, her pregnancy, her preferences and the advice of her clinician. We do not provide treatment recommendations, advice on any pregnancy or interpretation of an individual's tests. Women who have concerns about their health or pregnancy should speak to a health professional.

Frequently asked questions

Why are definitions important in obstetric reviews?

Because outcomes such as preterm birth and preeclampsia are defined differently across studies and years, and mixing definitions can mislead.

How should twins be handled?

Record the unit of analysis, avoid counting babies from the same pregnancy as independent, and use estimates adjusted for clustering.

What is the best outcome in fertility trials?

Live birth per woman randomized, with multiple pregnancy reported. Per-cycle and per-transfer rates are less informative.

How do I review drug safety in pregnancy?

Mostly with observational studies, using adjusted estimates, attention to timing of exposure and absolute risks, and reporting under MOOSE.

What is a core outcome set?

An agreed minimum set of outcomes that all trials in an area should report, which reduces variation and selective reporting.

Do you give pregnancy or treatment advice?

No. The service provides research and evidence-synthesis support only.

References

  1. Chalmers I, Enkin M, Keirse MJNC, editors. Effective care in pregnancy and childbirth. Oxford: Oxford University Press; 1989.
  2. Khan K. The CROWN initiative: journal editors invite researchers to develop core outcomes in women's health. BJOG. 2016;123(Suppl 3):103-104.
  3. Sterne JA, Hernan MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
  4. Stroup DF, Berlin JA, Morton SC, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-2012.
  5. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  6. WHO ACTION Trials Collaborators. Antenatal dexamethasone for early preterm birth in low-resource countries. N Engl J Med. 2020;383(26):2514-2525.
  7. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.