Method

Individual participant data meta-analysis

An individual participant data (IPD) meta-analysis collects the raw, participant-level data from each study instead of relying on published summaries. It is slower and costlier than a conventional review, and in return it allows analyses that aggregate data cannot support. This page explains when it is worth doing, how the analysis works and what can go wrong.

What an IPD meta-analysis is

A conventional meta-analysis combines summary results taken from publications: a mean difference, a risk ratio, a hazard ratio with its interval. An IPD meta-analysis obtains the underlying records of the participants in each study, who they were, what treatment they received and what happened to them, and analyses them directly. It is sometimes called the gold standard of evidence synthesis, a phrase that overstates the case but captures what the data allow. The review still begins with a systematic search and explicit criteria. The difference lies in what is collected from the eligible studies.

Most IPD meta-analyses are collaborative projects. The review team invites the investigators of each eligible study to share their data, and often to join a collaborative group that agrees the protocol, checks the analyses and shares authorship of the output. The Cochrane Handbook and the PRISMA-IPD statement describe the process.

What participant-level data make possible

Several advantages follow from having the raw data.

  • Consistent analysis. The same outcome definition, time point, population and statistical model can be applied in every study, which removes variation due to differing analytic choices in the publications.
  • Unpublished and unreported data. Outcomes and participants that were not in the publication can be included, and the data can be checked for errors and for selective reporting.
  • Time-to-event analysis. Hazard ratios can be estimated from the original follow-up times, with checks of the proportional hazards assumption, and restricted mean survival times can be calculated, which is difficult from published curves.
  • Participant-level subgroups and interactions. Whether a treatment works differently by age, sex, severity or a biomarker can be examined across all participants, with much greater power and validity than comparing study-level averages.
  • Handling missing data and adjustment. Confounders can be adjusted for in a standard way, and missing data handled consistently.
  • Longer follow-up and updated data. Investigators may provide follow-up beyond that published.
  • Prognostic and prediction modelling. IPD from many studies supports the development and validation of prediction models, and the examination of how performance varies by setting.

When it is worth the effort

An IPD project often takes several years and needs the cooperation of many investigators, with legal agreements and data management. It is justified when the question depends on something aggregate data cannot give. Typical cases are the study of treatment-covariate interactions to identify who benefits, a time-to-event outcome with poor reporting, a question for which published data are sparse or inconsistent, a need to examine outcomes at a standardized time point, the development of a prediction model, or a question of high importance for clinical practice or policy. It is less justified when the published data are complete and the aggregate method is enough for the question, which is the case for many reviews.

Feasibility depends on data availability. If only half of the eligible studies agree to share, the IPD set may be a selected sample, and availability bias becomes a concern. Compare the IPD studies with those that did not provide data, and where possible combine IPD with aggregate data from the others, using methods designed for this.

Conducting an IPD meta-analysis

  1. Protocol and registration. Specify the question, eligibility, the data items, the analysis plan including the interactions to be tested, and register it, for example in PROSPERO.
  2. Systematic search and identification. As for any systematic review, with attention to unpublished and ongoing studies, since an IPD collaboration can include them.
  3. Invitation and agreement. Contact investigators, explain the aims and the data needed, and agree terms of use, authorship and confidentiality. Data sharing platforms and repositories can simplify access for some trials.
  4. Data receipt and anonymization. Receive data in secure form, with a data dictionary, and ensure that privacy and ethics rules are met.
  5. Checking and harmonization. Check each dataset for completeness, internal consistency, plausible values, balance at baseline, the integrity of randomization and agreement with the publication. Recode variables to common definitions. Query problems with the investigators.
  6. Risk of bias. Assess each study, using the IPD to verify features such as randomization and attrition, which the publication may not describe.
  7. Analysis. Apply the prespecified model, as below.
  8. Interpretation and feedback. Share results with investigators for comment before publication, as is usual in collaborative projects.
  9. Reporting. Follow PRISMA-IPD, which extends PRISMA with items on data collection, checking and analysis.

Two-stage and one-stage analysis

There are two main ways of analysing IPD.

Two-stage. In the first stage, each study's data are analysed separately with the same model, producing an effect estimate and its variance, for example a log hazard ratio, or an interaction coefficient. In the second stage, those estimates are combined with a conventional meta-analytic model. This is simple, transparent, uses familiar tools and makes it easy to draw forest plots. It relies on large-sample approximations in the first stage, and can be awkward when studies are small or events are rare.

One-stage. All participants are analysed in a single model, typically a mixed-effects regression with a separate intercept (or baseline hazard) for each study and, usually, a random treatment effect. It is exact for binary and time-to-event outcomes, copes with sparse data, and allows complex models, such as non-linear covariate effects or multiple outcomes. It is more complex, and its assumptions about the distribution of random effects need care. Many authors recommend running both, since when the models are specified equivalently they usually agree, and a discrepancy points to a problem.

In both approaches, the study must be treated as a cluster. Pooling all participants as if from one trial breaks randomization, and can reverse the direction of an effect through a statistical effect called Simpson's paradox.

Treatment-covariate interactions and ecological bias

Asking whether a treatment effect varies by a participant characteristic is the most distinctive use of IPD, and also the one most open to error. Two quite different associations can be mixed up. The within-trial interaction compares treated and control participants of different ages inside each trial. The across-trial association compares the overall treatment effects of trials whose participants differ on average in age. The first is a participant-level quantity and is protected from confounding by randomization. The second is an ecological association, which can be distorted by anything that differs between trials.

A simulated example makes the point. Four trials report an interaction between treatment and age, estimated within each trial, and also differ in mean age and overall effect.

Four simulated trials: mean age, overall effect and within-trial interaction per year of age
TrialMean ageOverall effect (log scale)Within-trial interactionSE
Trial 145-0.10-0.0100.004
Trial 252-0.40-0.0120.005
Trial 360-0.35-0.0080.006
Trial 468-0.90-0.0150.007

Pooling the within-trial interactions gives -0.0108 per year (standard error 0.0026). A regression of overall effect on mean age across the trials gives a slope of -0.0305 per year, about 3 times larger. The across-trial slope is much steeper because trials with older participants also differ in other ways, such as setting, dose and follow-up. Using it would greatly overstate how much age matters. A good IPD analysis separates the two, by centering the covariate within each trial and modelling the trial mean separately, and bases conclusions on the within-trial estimate. The data are simulated.

Missing data, quality and availability bias

Having the data does not remove problems of quality. Checks should look for impossible values, duplicated participants, unbalanced baseline characteristics that suggest a failure of randomization, and discrepancies with the publication. If serious problems appear, they should be discussed with the investigators, and the study may be excluded or analysed with sensitivity analysis. Missing outcome data can be handled with multiple imputation, done within each study so as to respect the clustering, or with likelihood-based methods under stated assumptions.

Availability bias arises when the studies that supply data differ from those that do not, for example if investigators of trials with unfavorable results are less willing to share. Report how many eligible studies and participants were obtained, compare the IPD and non-IPD studies and, if aggregate data are available for the rest, combine both sources in a model that distinguishes them.

Reporting

PRISMA-IPD adds items to the PRISMA checklist for reporting how IPD were sought, which studies provided data, how data were checked, how the analyses handled the clustering and how the participant-level results were derived. Authors should state the number of studies and participants eligible and obtained, the reasons for missing data, the models, the handling of interactions and the sensitivity analyses. A flow diagram showing the progress from eligible to included studies and participants is expected. Data sharing agreements often limit what can be made public, and the data availability statement should explain the restrictions.

Limitations

IPD meta-analysis is resource-intensive, depends on the goodwill of investigators, and may deliver data on only a subset of studies. It requires data management and statistical skills, ethics and legal approvals, and months or years of work. Even with IPD, the conclusions are limited by the quality and comparability of the underlying studies. Subgroup analyses remain observational comparisons when the covariate was not randomized, and multiple interaction tests carry the risk of false positives, so those that matter should be prespecified and limited in number. The results describe populations of studies and do not give advice for individual patients.

IPD meta-analysis informs research and policy. It is not a substitute for clinical judgment about an individual patient.

How we can help

Support

Support for individual participant data

The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.

Get a quoteTell us your question, the number of studies and your target journal.

Frequently asked questions

What is an IPD meta-analysis?

A meta-analysis that uses the raw participant-level data from each eligible study, obtained from the investigators, instead of summary results from publications.

When should I choose IPD over aggregate data?

When the question needs participant-level analyses, such as treatment-covariate interactions or time-to-event outcomes, or when published data are incomplete or inconsistent. For many questions, aggregate data are enough.

What is the difference between one-stage and two-stage analysis?

Two-stage analyses each study first and then pools the estimates. One-stage fits one model to all participants with study-specific terms. Well-specified versions usually agree.

What is ecological bias in IPD?

A distortion that arises when an association between trial-level averages is read as if it applied to individuals. Separate within-trial from across-trial effects and base conclusions on the within-trial effect.

What if some investigators will not share data?

Report the proportion obtained, compare the IPD and non-IPD studies and, if aggregate data exist for the others, combine them in a model that accounts for the different sources.

Which reporting guideline applies?

PRISMA-IPD, an extension of PRISMA, with items on obtaining, checking and analysing individual participant data.

References

  1. Stewart LA, Clarke M, Rovers M, et al. Preferred Reporting Items for Systematic Review and Meta-Analyses of individual participant data: the PRISMA-IPD Statement. JAMA. 2015;313(16):1657-1665.
  2. Riley RD, Lambert PC, Abo-Zaid G. Meta-analysis of individual participant data: rationale, conduct, and reporting. BMJ. 2010;340:c221.
  3. Riley RD, Tierney JF, Stewart LA, eds. Individual Participant Data Meta-Analysis: A Handbook for Healthcare Research. Wiley; 2021.
  4. Fisher DJ, Copas AJ, Tierney JF, Parmar MKB. A critical review of methods for the assessment of patient-level interactions in individual participant data meta-analysis of randomized trials, and guidance for practitioners. J Clin Epidemiol. 2011;64(9):949-967.
  5. Berlin JA, Santanna J, Schmid CH, Szczech LA, Feldman HI. Individual patient- versus group-level data meta-regressions for the investigation of treatment effect modifiers: ecological bias rears its ugly head. Stat Med. 2002;21(3):371-387.
  6. Debray TPA, Moons KGM, van Valkenhoef G, et al. Get real in individual participant data (IPD) meta-analysis: a review of the methodology. Res Synth Methods. 2015;6(4):293-309.
  7. Tierney JF, Vale C, Riley R, et al. Individual participant data (IPD) meta-analyses of randomised controlled trials: guidance on their use. PLoS Med. 2015;12(7):e1001855.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.