What incidence is and how it differs from prevalence
Incidence counts new cases arising in a population during a period. Prevalence counts existing cases at one time or over a period. Incidence measures the rate at which people become ill, and prevalence measures the burden at a point. They are related, since prevalence is approximately incidence multiplied by the average duration of the condition, but a review of one cannot answer questions about the other.
Two forms of incidence are used. The incidence rate, also called incidence density, is the number of new events divided by the person-time at risk, for example 3 cases per 1,000 person-years. It handles different follow-up lengths naturally. The cumulative incidence or risk is the proportion of a fixed group that develops the outcome over a fixed period, for example the 5-year risk of a disease. It needs complete follow-up or a method for censored data. Reviews should say which is being pooled, because they are different quantities and cannot be combined.
Pooling incidence rates
The usual approach works on the log scale. For a study with e events and person-time T, the incidence rate is e / T, the log rate is its natural logarithm and the standard error of the log rate is approximately 1 / sqrt(e), since the number of events is Poisson-distributed. Notice that precision depends on the number of events, not on the number of people, which is why a large cohort with few events has a wide interval. The studies are pooled by inverse-variance weighting, and the result is exponentiated.
| Study | Events | Person-years | Rate per 1,000 | 95% CI |
|---|---|---|---|---|
| Cohort A | 18 | 12,000 | 1.50 | 0.95 to 2.38 |
| Cohort B | 42 | 25,000 | 1.68 | 1.24 to 2.27 |
| Cohort C | 9 | 8,000 | 1.12 | 0.59 to 2.16 |
| Cohort D | 65 | 31,000 | 2.10 | 1.64 to 2.67 |
| Cohort E | 27 | 9,000 | 3.00 | 2.06 to 4.37 |
The fixed-effect pooled rate is 1.96 per 1,000 person-years, but Q is 10.2 on 4 degrees of freedom and I-squared is about 61 percent, so the studies do not share one rate. The random-effects pooled rate is 1.88 per 1,000 person-years (95 percent CI 1.44 to 2.46), and the 95 percent prediction interval, with a t distribution on 3 degrees of freedom, is 0.80 to 4.43. The prediction interval is far wider than the confidence interval, which is realistic: the rate in a new population could differ widely. The cohorts are simulated.
It is the prediction interval, more than the pooled mean, that tells a reader what to expect elsewhere. Incidence rates vary with age, sex, geography, time period and the definition and ascertainment of cases, so a single pooled number rarely describes any actual population.
Poisson and binomial mixed models
The log-rate approach is an approximation. It fails when counts are small or zero, because the standard error depends on the observed count. A better approach for incidence rates is a Poisson mixed model, which treats the events in each study as Poisson with mean equal to the rate times person-time, and the log rate as a random effect across studies. It uses the exact likelihood, handles zero counts without correction, and allows the inclusion of covariates and of studies with different person-time. For cumulative incidence, the corresponding model is a binomial mixed model with a logit or log link. Both are available in R (for example with metafor and lme4) and in Stata.
Some authors use transformations, such as the Freeman-Tukey double arcsine or the square root, to stabilize the variance. Evidence from simulation and theory suggests these can give misleading results, in particular when back-transformed, so mixed models or the log transformation with exact methods are generally preferred. Whichever method is used, the sensitivity of the result to the choice should be checked.
Zero events
Studies in which no event was observed pose the same problem as in other meta-analyses of sparse data. The log of zero is undefined, and the standard error 1 / sqrt(0) is infinite. Adding 0.5 to the count is common and arbitrary, and it biases the pooled rate upward when many studies have no events. Better options include the Poisson mixed model, the exact Poisson confidence interval for each study, which is defined even for zero events, and Bayesian models. Report the number of studies with zero events and show the sensitivity of the pooled rate to the method. A study with zero events in a short period does not show that the rate is zero. The upper limit of an exact interval is about 3 divided by the person-time, for a 95 percent interval, by the rule of three.
Incidence rate ratios and comparisons
When studies compare groups, such as exposed and unexposed, the effect measure is the incidence rate ratio, the ratio of the two rates. It is pooled on the log scale with variance equal to the sum of the reciprocals of the event counts in the two groups. Hazard ratios from time-to-event analyses estimate a related quantity but not an identical one, and the two should not be mixed in a single pool without thought. The same cautions apply as for risk ratios: with adjusted estimates from observational studies, pool the adjusted ratios, record the covariates and consider confounding. For comparing incidence across populations or time, standardization by age and sex matters, because crude rates differ simply by the age structure of the populations.
Sources of heterogeneity and bias
- Case definition and ascertainment. Active surveillance finds more cases than passive reporting. Registry-based studies and clinical cohorts differ in completeness.
- Population at risk. The denominator must exclude people with prevalent disease. Studies that fail to do so give inflated rates.
- Time period and follow-up. Rates change over time, and the person-time must be measured consistently. Short follow-up may capture early events only.
- Age and sex. Rates differ strongly by age, so report age-specific or standardized estimates where possible.
- Geography and health systems. Access to diagnosis varies.
- Selection. Hospital-based or trial populations may not represent the community.
- Competing risks. People who die of another cause no longer contribute risk, which affects cumulative incidence.
Meta-regression can explore some of these, with limits set by the number of studies. Risk of bias tools for incidence and prevalence studies, such as the Joanna Briggs Institute checklist for prevalence studies and the tool by Hoy and colleagues, can be adapted. The MOOSE guideline for observational reviews applies, along with PRISMA.
Conducting an incidence review
- Define the incidence measure, the population, the time period and the case definition, in a protocol.
- Search for cohort studies, registries and surveillance reports, including grey literature, which is a large source of incidence data.
- Extract events, person-time or numbers followed, follow-up duration, age and sex distribution and case ascertainment.
- Assess risk of bias with a suitable tool.
- Choose the model and handle zero counts, in a prespecified way.
- Pool, describe heterogeneity with a prediction interval, and run meta-regression or subgroup analyses by age, region and period.
- Report to PRISMA 2020 and MOOSE, with the rate scale clearly labeled.
Presenting incidence results
Present each study's rate with its interval in a forest plot on a log axis, label the unit clearly (per 1,000 person-years, per 100,000 population per year) and use the same unit throughout. Alongside the pooled rate, show the prediction interval and the heterogeneity statistics. When studies cover different age groups, show age-specific pooled rates in a table or a plot, since a single overall rate is hard to interpret. Maps and plots by period are helpful when geography or time matter. State the population to which the pooled rate applies and the period it covers, and warn against applying it to other populations without adjustment. If cumulative incidence at several time points is pooled, say whether the same studies contribute at each time point, because the composition changes with follow-up. For policy readers, converting rates into expected numbers of cases in a population of a given size, with the uncertainty, is often more useful than the rate itself, provided the assumptions behind the conversion are stated and the uncertainty is carried through to the final figure.
Limitations
Pooled incidence is often very heterogeneous, and the average may not describe any real population. Different case definitions and ascertainment mean the studies may not measure the same thing. Rates from different time periods are not comparable when incidence changes. The methods rest on assumptions, such as Poisson counts and independence, which can fail with overdispersion or clustering. A pooled rate is a summary of studies and not an estimate of the rate in a specific place. Results are for research and planning, and do not give advice about an individual's risk, and the pooled figures should be read together with the quality of the underlying surveillance.
Incidence estimates describe populations. They do not predict whether an individual person will develop a condition.
How we can help
Support for incidence meta-analysis
The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.
Feasibility check
A review of your studies and data to confirm that the method is suitable and which approach fits.
Analysis and figures
The analysis run to a prespecified plan, with forest plots and the other figures.
Methods and results text
Written for the manuscript and aligned with PRISMA 2020 or the relevant extension.
Manuscript and submission
Optional: the full paper, the reporting checklist and the submission materials.
Frequently asked questions
What is the difference between incidence and prevalence?
Incidence counts new cases over time and measures the rate of occurrence. Prevalence counts existing cases and measures burden. Reviews of one cannot answer questions about the other.
What is the difference between incidence rate and cumulative incidence?
The incidence rate is events per unit of person-time. Cumulative incidence is the proportion of a group developing the outcome in a fixed period. They are different quantities and are pooled separately.
How do I pool incidence rates?
Use the log of each rate with variance of one over the number of events, with inverse-variance weights and a random-effects model, or fit a Poisson mixed model, which handles small counts better.
What do I do with studies that observed no events?
Use exact or model-based methods, such as a Poisson mixed model or exact intervals, and avoid ad hoc corrections. Report how many studies had zero events.
Why is a prediction interval important?
Incidence usually varies greatly between populations, so the prediction interval shows the range expected in a new setting, which the confidence interval for the average does not.
Can I combine hazard ratios and rate ratios?
Not without thought. They estimate related but different quantities. Analyse them separately or justify combining them.
References
- Rothman KJ, Greenland S, Lash TL. Modern Epidemiology. 3rd ed. Lippincott Williams and Wilkins; 2008.
- Stijnen T, Hamza TH, Ozdemir P. Random effects meta-analysis of event outcome in the framework of the generalized linear mixed model with applications in sparse data. Stat Med. 2010;29(29):3046-3067.
- Schwarzer G, Chemaitelly H, Abu-Raddad LJ, Rucker G. Seriously misleading results using inverse of Freeman-Tukey double arcsine transformation in meta-analysis of single proportions. Res Synth Methods. 2019;10(3):476-483.
- Hoy D, Brooks P, Woolf A, et al. Assessing risk of bias in prevalence studies: modification of an existing tool and evidence of interrater agreement. J Clin Epidemiol. 2012;65(9):934-939.
- Munn Z, Moola S, Lisy K, Riitano D, Tufanaru C. Methodological guidance for systematic reviews of observational epidemiological studies reporting prevalence and cumulative incidence data. Int J Evid Based Healthc. 2015;13(3):147-153.
- Stroup DF, Berlin JA, Morton SC, et al. Meta-analysis of observational studies in epidemiology: a proposal for reporting. JAMA. 2000;283(15):2008-2012.
- Hanley JA, Lippman-Hand A. If nothing goes wrong, is everything all right? Interpreting zero numerators. JAMA. 1983;249(13):1743-1745.