Evidence synthesis in pediatrics
Children are not small adults. Their physiology, drug handling, risk of disease and ability to report symptoms change with age, and the same intervention can have different effects in a newborn, a toddler and a teenager. Neonatal medicine has an especially strong tradition of meta-analysis: the Cochrane Neonatal group and trial networks have produced reviews of respiratory support, nutrition, infection control and neuroprotection for preterm infants, and some of these reviews established practice that is now routine. Other areas, such as childhood asthma, infections, mental health and obesity, rely on trials that are smaller and more varied.
Research with children faces ethical and practical limits. Recruitment is slower, consent is given by parents, placebo or no-treatment controls are harder to justify, and follow-up to adulthood is expensive. As a result, evidence is often thin, trials are underpowered, and many treatments are used off-label on the basis of adult data. A review should state this plainly and avoid concluding that a treatment works because no trial has shown that it does not.
The methods follow meta-analysis and systematic review, with adjustments for age and development.
Age, development and effect modification
Age is a major source of heterogeneity. Reviews should define the age range in the question and avoid pooling infants with adolescents unless there is a reason to expect similar effects. Where studies cover a wide span, subgroup analysis or meta-regression on mean age can examine variation, with the limits of study-level data. In neonatal research, gestational age and birth weight define subgroups: extremely preterm infants have very different risks and treatment responses from late preterm infants. Reviews should extract the gestational age range and consider separate analyses.
Growth and developmental outcomes are measured as z-scores or percentiles relative to reference charts. Reviews should record which reference was used (for example, World Health Organization standards or national charts), because z-scores differ between references. Standardized mean differences for developmental scores depend on the instrument and the age at testing, and scores obtained at 2 years often predict later outcomes poorly. When long-term follow-up is available, it is worth extracting and reporting it, with the proportion of the original sample that was followed.
Neonatal outcomes and composite endpoints
Neonatal trials often use composite outcomes, such as death or bronchopulmonary dysplasia, death or neurodevelopmental impairment, and death or severe brain injury. Death is a competing risk: an intervention that reduces mortality may appear to increase morbidity in survivors because more fragile infants survive. Reporting death and morbidity among survivors separately as well as in the composite helps the reader. Definitions vary, for example the definition of bronchopulmonary dysplasia at 28 days or at 36 weeks of postmenstrual age, and a review should tabulate them and use consistent definitions, or analyze groups of definitions separately.
Neurodevelopmental impairment is assessed at 18 to 24 months with instruments such as the Bayley Scales, and a substantial share of infants are lost to follow-up. Attrition may be related to outcome, since families of children with problems may be harder or easier to reach, so reviews should consider the completeness of follow-up when rating certainty.
Cluster trials, community interventions and small studies
Many pediatric interventions are delivered in schools, clinics or communities: vaccination campaigns, nutrition programs, hygiene, parenting support. They are often cluster-randomized, and the analysis must take clustering into account, using an adjustment based on the intraclass correlation when authors did not. Treating children as independent when they are clustered makes confidence intervals too narrow. Reviews should record the unit of randomization, the number of clusters and whether the analysis adjusted for them.
In rare pediatric diseases, trials can have a few dozen participants. A meta-analysis with few studies has imprecise estimates of heterogeneity, and a Hartung-Knapp adjustment, a prediction interval and, where justified, a Bayesian analysis with an informative prior from adult data can help. Using adult data as prior information should be explicit, with a sensitivity analysis on the strength of borrowing, because children may respond differently.
Safety, adverse events and off-label use
Adverse events in children may differ from those in adults and may show up only in follow-up. Drug trials in children are often too small to detect rare events, and meta-analysis of adverse events is subject to sparse-data problems: zero events in one or both arms are common. Methods such as Mantel-Haenszel with or without continuity corrections, the Peto odds ratio for balanced designs, and beta-binomial models are available, and the choice should be stated. The guide on zero-event studies explains these choices. Observational data from registries and pharmacovigilance systems provide supporting information about rare harms, with confounding and reporting biases to consider.
Vaccine safety and effectiveness reviews follow the same principles, with attention to the age at vaccination, schedule, and case definitions.
Diagnostic and prognostic questions
Pediatric diagnostic reviews, for example of tests for urinary tract infection, appendicitis or serious bacterial infection in febrile infants, use bivariate or hierarchical models for sensitivity and specificity. Spectrum bias is a concern, since studies in specialist centers include sicker children than primary care. Prognostic reviews of outcomes such as asthma persistence or neurodevelopment need attention to the age at which prediction is made and to the timing of follow-up. See diagnostic accuracy meta-analysis and prognostic meta-analysis.
A worked reading of a neonatal result
Suppose a review pools trials of a respiratory strategy for preterm infants, with the outcome death or bronchopulmonary dysplasia. In the control groups 30% of infants had the outcome, and in the intervention groups 24% did. The numbers are invented to show the reading. The risk ratio is 0.24/0.30 = 0.80, the absolute risk reduction is 6 percentage points, and the number needed to treat is 1/0.06, which is about 17.
A careful reader asks four questions. First, how much of the composite is death and how much is morbidity in survivors, and do both move in the same direction? Second, what is the gestational age range in the trials, and does the effect differ for the most immature infants? Third, how many infants were followed to the age at which the outcome was assessed? Fourth, what is the certainty of the evidence: did the trials mask the caregivers, and are the confidence limits compatible with a small or no benefit? A summary that gives only the risk ratio leaves these questions open. A good review reports the absolute effect, the components, the age range and the certainty together.
Families, equity and context
Parents and caregivers are partners in pediatric research, and their priorities may differ from those chosen by researchers. A review that includes outcomes such as family stress, time in hospital, feeding and school attendance, alongside clinical events, reflects what matters in practice. Qualitative evidence on parents' experiences can be added through a mixed-methods or qualitative synthesis, see mixed-methods review.
Child health varies with poverty, region, ethnicity and access to care. Most trials are done in high-income countries, while most child deaths occur in low- and middle-income countries where health systems and baseline risks differ. The review should record where trials were done and consider whether results apply to other settings, using the PRISMA-Equity extension as a guide. Interventions such as zinc, oral rehydration or vaccines have large effects in settings with high baseline risk, and small absolute effects where risk is low, so absolute effects should be calculated for realistic baseline risks in the setting of interest, using local surveillance data or the control-group risks of trials done in comparable settings. Where the baseline risk is uncertain, presenting a low, medium and high scenario is clearer than choosing one number. Readers can then match their own setting to the nearest scenario.
Common pitfalls we look for
- Pooling children of very different ages or gestational ages as if they responded alike.
- Using composite outcomes without separating death from morbidity in survivors.
- Ignoring clustering in school or community trials.
- Assuming adult results apply to children without pediatric data.
- Not reporting loss to follow-up in long-term outcome studies.
- Mixing growth references when pooling z-scores.
Planning and reporting
The protocol defines the age range, the setting, the intervention and comparator, the outcomes with definitions and the time of assessment, and the planned subgroup analyses by age or gestation. Searches cover MEDLINE, Embase, CENTRAL, CINAHL and trial registries, with LILACS or regional databases for global child health. Risk of bias is assessed with RoB 2 for trials, with cluster-trial additions, and ROBINS-I for non-randomized studies. Certainty is rated with GRADE, with attention to imprecision and indirectness if evidence comes from adults or from different age groups. Reporting follows PRISMA 2020, and a protocol is registered in PROSPERO.
Families and clinicians can help to choose outcomes that matter. Core outcome sets, available in some neonatal and pediatric areas, support consistency across trials and reviews.
How we support research projects in this area
From a clinical question to a published review
Support can cover a whole review or a single stage. The scope is agreed at the start.
Question and protocol
A structured question, eligibility criteria and an analysis plan, with registration prepared where appropriate.
Searching and extraction
Search strategies for the relevant databases and registries, screening and data extraction, and risk-of-bias assessment by design.
Synthesis
Pairwise, network, diagnostic accuracy, prognostic or dose-response analysis, with a GRADE assessment for each outcome.
Manuscript and submission
Reporting-guideline checklists, the manuscript and the preparation of submission materials.
Boundaries of this service
A review of pediatric studies describes effects in groups of children. It does not tell parents or clinicians what to do for an individual child, whose care depends on age, weight, other conditions and the advice of the child's doctor. We do not provide treatment recommendations, dosing advice or interpretation of an individual child's results. For dosing of a child's medicine, the product label, the child's doctor or a pharmacist is the right source.
Frequently asked questions
Can I pool trials across different ages of children?
Only if effects are expected to be similar. Otherwise analyze age groups separately or examine age as a moderator.
How should I treat composite neonatal outcomes?
Report death, morbidity in survivors and the composite separately, and tabulate the definitions used.
How do I handle cluster-randomized trials?
Use effect estimates adjusted for clustering, or adjust standard errors using the intraclass correlation and test the assumption.
What if very few pediatric trials exist?
Say so. A Bayesian analysis using adult data as a prior can be considered, with the degree of borrowing stated and varied.
How should rare adverse events be analyzed?
With methods for sparse data such as Mantel-Haenszel, Peto or beta-binomial models, stating the handling of zero events.
Do you give dosing or treatment advice for a child?
No. The service provides research and evidence-synthesis support only.
References
- Higgins JPT, Thomas J, Chandler J, et al., editors. Cochrane Handbook for Systematic Reviews of Interventions. Cochrane; current version available at training.cochrane.org/handbook.
- Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
- Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
- Bradburn MJ, Deeks JJ, Berlin JA, Russell Localio A. Much ado about nothing: a comparison of the performance of meta-analytical methods with rare events. Stat Med. 2007;26(1):53-77.
- Hedges LV. Effect sizes in cluster-randomized designs. J Educ Behav Stat. 2007;32(4):341-370.
- Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.