Meta-analysis in public health and epidemiology

Public health and epidemiology study how common diseases are, what causes them and which population interventions work. Most evidence is observational and comes from surveys, cohorts and routine data. Reviews must deal with differences in populations and measurement, confounding and the complexity of interventions delivered to communities.

Evidence synthesis in public health and epidemiology

Public health uses research synthesis to learn about disease burden, risk factors and interventions at the level of populations. Examples include pooled estimates of the prevalence of hypertension or diabetes across countries, meta-analyses of smoking and cancer, reviews of tax and labeling policies on diet, and syntheses of programs for vaccination uptake or injury prevention. The Global Burden of Disease project combines many data sources with statistical models and is a related but different endeavor from a systematic review.

Several conditions make this a demanding field. Populations differ in age structure, ethnicity, income and health systems, so heterogeneity is large and expected. Studies use different case definitions, diagnostic methods and sampling frames. Most designs are observational, with confounding. Interventions are complex and delivered to groups, often evaluated by natural experiments where randomization is impossible. And policy questions require attention to equity: whether interventions narrow or widen differences among groups.

Methods follow meta-analysis and systematic review, with the specific points below.

Prevalence and incidence synthesis

Pooled prevalence is a common use of meta-analysis in public health. Each study contributes a proportion, a transformation (logit or Freeman-Tukey double arcsine, although the latter has known problems on back-transformation) puts it on a scale suited to modelling, and a random-effects model combines them. For example, a survey finding 18 cases among 150 people has a prevalence of 0.12. On the logit scale the estimate is -1.992 with a standard error of 0.251, which back-transforms to a 95 percent interval of 0.077 to 0.182. The numbers are invented. A generalized linear mixed model with a binomial likelihood avoids the transformation and handles proportions near 0 or 1 better. The guide to prevalence meta-analysis sets out the options.

Heterogeneity is generally very high, often with I-squared above 90 percent, because studies differ in time, place, age and method. A single pooled prevalence from such studies has little meaning, and the prediction interval and meta-regression on age, region, sampling method and diagnostic method tell more. Incidence is analyzed as rates with person-time, see incidence meta-analysis. Reviews should specify the denominators: a prevalence among those tested is not a prevalence in the population.

Observational studies and confounding

Reviews of risk factors combine cohort and case-control studies. The MOOSE reporting guideline asks for details of the search, the selection, the assessment of exposure and outcome, the adjustment for confounders and the handling of heterogeneity. Pooling adjusted estimates is preferred over crude estimates, but adjustment sets differ across studies, which is a source of heterogeneity. A review should compare the estimates adjusted for few and many confounders and examine the E-value or similar measures of how strong unmeasured confounding would need to be to explain the result. Case-control studies sample on the outcome, and the odds ratio approximates the risk ratio only when the outcome is rare. Cohort studies suffer from loss to follow-up and from changes in exposure over time.

Reverse causation, in which early disease changes exposure, such as weight loss before a cancer diagnosis, can reverse the apparent direction of an association. Analyses that exclude the first years of follow-up address this. Risk-of-bias assessment uses ROBINS-E or other tools for exposure studies, and the certainty of evidence begins low in GRADE and may be upgraded for large effects or dose-response gradients.

Population interventions and natural experiments

Taxes on sugary drinks, smoke-free laws, school meal reforms, housing programs and vaccination campaigns cannot usually be randomized in the same way as drugs. Evaluations use controlled before-after studies, interrupted time series, difference-in-differences and synthetic control methods. Each rests on assumptions about the absence of other changes at the time of the intervention. A review should record the design and assess the plausibility of the assumptions: parallel trends, no co-interventions, enough data points. Effects reported as changes in level or slope must be converted to a common metric before pooling, and results from different designs are often better presented side by side.

Policy effects often differ by place and time. Reviews can use meta-regression to explain the variation by tax rate, population or implementation, and qualitative evidence to understand mechanisms. Where the effect on health outcomes needs many years, reviews may rely on intermediate outcomes like purchases or consumption, with the link to health stated as an assumption.

Equity and social determinants

Public health aims to reduce health inequalities. A review that reports only an average effect may miss a pattern in which an intervention helps advantaged groups more than disadvantaged ones, as some information-based campaigns do. The PROGRESS-Plus framework lists characteristics along which inequality is described: place of residence, race or ethnicity, occupation, gender, religion, education, socioeconomic status and social capital. Reviews should state whether trials reported results by these factors and analyze effects by them when possible. The PRISMA-Equity extension guides reporting. Many studies do not report such analyses, which is itself a finding. Small subgroups give imprecise estimates, and the interaction test is more informative than separate subgroup results.

Surveillance, routine data and global estimates

Routine data from registries, vital statistics and surveillance systems are not studies in the usual sense, and the same cases may appear in several reports. Reviews that use them need to identify overlap and decide on the primary source. Quality varies between countries, with underreporting in settings where systems are weak. Modeling studies that estimate burden by combining sources make assumptions that should be listed. Reviews should keep estimates from models separate from those based on direct measurement, and present the data sources used for each.

Screening, vaccination programs and health services research

Screening programs raise their own issues. The benefit of screening is judged by reduced disease-specific mortality in randomized trials, and the harms include false positives, overdiagnosis and overtreatment. Lead-time bias and length bias make survival from diagnosis a poor measure of benefit. Reviews of screening should report mortality and harms, the number needed to screen and, where possible, overall mortality, and should be careful about all-cause mortality being too insensitive to detect the benefit for a single disease. Observational studies of screening are open to self-selection bias, since people who attend are healthier than those who do not.

Health services research covers access, quality and cost: for example, the effect of insurance coverage on use, hospital quality improvement programs, community health workers and telehealth. Randomized evidence exists for some of these, such as community health worker trials, but much comes from observational analyses of administrative data. Reviews should specify the system context, since results from one health system do not transfer easily to another, and should record whether authors accounted for clustering of patients within hospitals.

Environmental health and global health

Environmental exposures such as air pollution, heat, lead and water quality are studied in time-series, cohort and case-crossover designs. Exposure assessment is a major source of error, as exposure estimates come from monitors or models rather than personal measures, and the concentration-response function is often non-linear. Pooled estimates per unit of exposure, such as per 10 micrograms per cubic meter of fine particulate matter, are used in burden assessment, and the choice of function at low and high concentrations affects results. Reviews should show the range of exposures in the data and avoid extrapolation.

Global health reviews of programs in low- and middle-income countries must consider how studies are located and reported, since databases index regional journals unevenly. Searching regional databases and grey literature helps, and reporting the countries represented lets readers see where evidence is thin. Disparities in research funding mean that conditions with the highest burden may have the least research. A review should therefore report the geographic distribution of its included studies, compare it with the distribution of the burden, and say which populations are missing. This helps policy makers judge whether findings from the available studies can guide action in their own setting, and where new research would be most valuable. Funders and public health agencies can use such a map of evidence to plan the next studies, and a living or regularly updated review can keep the map current as new surveys, cohorts and evaluations are published in the years that follow.

Common pitfalls we look for

  • Presenting a pooled prevalence from very different populations as a single figure.
  • Using the double arcsine back-transformation without checking.
  • Pooling crude and adjusted estimates.
  • Counting the same cohort or survey in several publications.
  • Ignoring reverse causation in exposure studies.
  • Reporting average effects only when equity is a stated aim.

Planning and reporting

The protocol states the population, the exposure or intervention, the outcome definitions, the designs and the method for pooling and exploring heterogeneity, including the prediction interval and meta-regression variables. Searches cover MEDLINE, Embase, Global Health, CINAHL, Scopus, regional databases and the websites of public health agencies. Reporting follows PRISMA 2020 and MOOSE, and for prevalence reviews the JBI critical appraisal checklist for prevalence studies is commonly used. A protocol can be registered in PROSPERO where eligible.

How we support research projects in this area

Support

From a clinical question to a published review

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A structured question, eligibility criteria and an analysis plan, with registration prepared where appropriate.

  • Searching and extraction

    Search strategies for the relevant databases and registries, screening and data extraction, and risk-of-bias assessment by design.

  • Synthesis

    Pairwise, network, diagnostic accuracy, prognostic or dose-response analysis, with a GRADE assessment for each outcome.

  • Manuscript and submission

    Reporting-guideline checklists, the manuscript and the preparation of submission materials.

Get a quoteDescribe your question, study types and target journal.

Boundaries of this service

A review of public health and epidemiological studies describes populations and groups. It does not provide clinical advice, public-health orders or policy decisions, and it does not tell any person what to do about their health. Policy decisions weigh evidence along with costs, values and feasibility, and belong to those responsible for them.

Frequently asked questions

Why is heterogeneity so high in prevalence reviews?

Because studies differ in population, time, age, sampling and diagnostic method. Prediction intervals and meta-regression are more informative than a single pooled value.

Which transformation should I use for proportions?

The logit is a common choice, and generalized linear mixed models are preferred in many cases. The double arcsine back-transformation can mislead.

How should I assess confounding in observational reviews?

Compare estimates with different adjustment sets, use tools such as ROBINS-E and consider measures like the E-value.

How are natural experiments combined?

By converting effects to a common metric, noting the design and assumptions, and often presenting different designs separately.

What is PROGRESS-Plus?

A framework of characteristics that shape health inequity, used to plan and report equity analyses.

Do you give public-health or clinical advice?

No. The service provides research and evidence-synthesis support only.

References

  1. Stroup DF, Berlin JA, Morton SC, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-2012.
  2. Welch V, Petticrew M, Tugwell P, et al. PRISMA-Equity 2012 extension: reporting guidelines for systematic reviews with a focus on health equity. PLoS Med. 2012;9(10):e1001333.
  3. Schwarzer G, Chemaitelly H, Abu-Raddad LJ, Rucker G. Seriously misleading results using inverse of Freeman-Tukey double arcsine transformation in meta-analysis of single proportions. Res Synth Methods. 2019;10(3):476-483.
  4. Munn Z, Moola S, Lisy K, Riitano D, Tufanaru C. Methodological guidance for systematic reviews of observational epidemiological studies reporting prevalence and cumulative incidence data. Int J Evid Based Healthc. 2015;13(3):147-153.
  5. VanderWeele TJ, Ding P. Sensitivity analysis in observational research: introducing the E-value. Ann Intern Med. 2017;167(4):268-274.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.