Method

Bayesian meta-analysis: the method

Bayesian meta-analysis expresses the same hierarchical model as random-effects meta-analysis but treats every unknown quantity as having a probability distribution. This page explains the model, the role of priors, how the model is fitted and checked, and how results are read and reported.

Bayesian inference in brief

In Bayesian statistics, uncertainty about an unknown quantity is expressed as a probability distribution. Before seeing the data, this is the prior distribution. The data enter through the likelihood, which gives the probability of the observed data for each possible value of the unknown. Bayes' theorem combines the two: the posterior distribution is proportional to the prior times the likelihood. The posterior is the complete answer, from which any summary can be drawn: a median, a mean, an interval that contains the unknown with stated probability, or the probability that it exceeds a threshold.

The difference from conventional meta-analysis lies in interpretation and in the use of prior information. A 95 percent credible interval is an interval that, given the model and the priors, contains the true value with probability 95 percent, which is the interpretation many people mistakenly give to a confidence interval. And prior information can be included explicitly when it is available and defensible. In practice many Bayesian meta-analyses use priors that carry little information, so that the results resemble those of a conventional analysis, with added benefits in how uncertainty is propagated and described. The Bayesian meta-analysis service describes the practical side.

The hierarchical model

The standard Bayesian model for meta-analysis has three levels. At the first, each study's observed estimate y_i is assumed to be normally distributed around the study's true effect theta_i with a variance v_i that is treated as known. At the second, the true effects are exchangeable draws from a normal distribution with mean mu and standard deviation tau. At the third, mu and tau have prior distributions. The overall mean mu is the quantity of main interest, tau describes the heterogeneity, and each theta_i is a study-specific effect that is partly shrunk toward the overall mean.

The shrinkage is one of the useful features. The posterior for each study's true effect combines its own data with what the other studies imply, so the estimates for small, imprecise studies are pulled more toward the mean than those of large, precise ones. This is the same phenomenon as in any hierarchical model and is often a better estimate of each study's effect than the study's own raw estimate. The model can be extended to binary data analyzed on the count scale, with a binomial likelihood and a logit link, and to meta-regression, multivariate and network structures.

Priors

Choosing the priors is the step that most distinguishes a Bayesian analysis, and it should be treated as part of the model specification and reported in full. Priors are of three broad types. Non-informative or vague priors aim to let the data dominate, but they are never quite neutral and can be unintentionally informative, especially for the variance parameters. Weakly informative priors rule out implausible values while leaving plausible ones open, and they are the recommended default in much current practice. Informative priors encode real information, from earlier studies, from related populations or from expert judgment, and need a justification that readers can assess.

For the overall effect mu, a normal prior centered at no effect with a wide variance is typical, and for a ratio measure it can be chosen so that implausibly large effects are excluded. For the heterogeneity standard deviation tau, the choice matters more, because it is poorly informed by the data when studies are few. Weakly informative half-normal or half-Cauchy priors on tau are often used, and empirical distributions based on large collections of published meta-analyses can inform priors for specific types of outcome and comparison. Very wide uniform priors on the variance or conventional inverse-gamma priors can distort the results with few studies and should be avoided or tested. See choosing priors.

Fitting the model

The posterior distributions of a meta-analysis model rarely have a closed form, so they are approximated by simulation. Markov chain Monte Carlo methods generate a long series of draws from the joint posterior, and the summaries are computed from the draws. Several chains are started from different points, and the first part of each is discarded as a warm-up. General-purpose software, including Stan and JAGS and packages that wrap them, makes the fitting routine. Some specialized packages for the normal-normal model use numerical integration and avoid simulation altogether, with results that can be computed quickly and exactly to numerical precision.

The simulations must be checked. Standard diagnostics are trace plots, to see whether the chains mix and settle; the potential scale reduction factor, which compares variation within and between chains and should be close to one; the effective sample size, which indicates how much independent information the draws contain; and, for Hamiltonian Monte Carlo, the number of divergent transitions, which signal regions of the posterior that the sampler cannot explore reliably. A model that has not converged gives numbers that look reasonable and may be wrong, so the diagnostics are reported along with the results.

Posterior summaries

The posterior distribution of the overall mean is summarized with a point estimate, usually the median or mean, and a credible interval, typically the central 95 percent interval or the highest-density interval. For heterogeneity the posterior of tau shows how much the data can say about it, and the interval is often wide. The posterior predictive distribution of the effect in a new study corresponds to a prediction interval in a conventional analysis, and incorporates the uncertainty in both mu and tau without the approximation that conventional prediction intervals require, which makes it especially useful.

Probability statements are the most distinctive output. The posterior probability that the effect is beneficial, that it exceeds a clinically meaningful threshold, or that one treatment ranks above another can be read directly. They are informative but depend on the prior, so they are reported with the prior sensitivity analysis. A statement such as there being a 95 percent probability of benefit has to be accompanied by the model and priors that produced it.

Sensitivity to priors and model

The conclusions of a Bayesian analysis are checked by repeating the analysis with alternative priors and, where relevant, alternative model structures. The priors for tau are varied most, since they have the greatest influence with few studies, and priors for the mean are varied to include skeptical and enthusiastic choices. The analyst reports whether the conclusions are stable. If they are, the data are doing the work. If not, the report states plainly that the conclusion depends on the prior and shows the range, which is more informative than a single number. Posterior predictive checks, which compare simulated data from the fitted model with the observed data, can reveal a poor fit of the model itself.

An illustration of shrinkage

Shrinkage is easiest to see with numbers. Suppose the overall mean effect is estimated at 0.30 and a small study reports 0.80 with a large standard error of 0.40, while a large study reports 0.25 with a standard error of 0.05. If the between-study standard deviation is 0.10, the small study's estimate is pulled strongly toward the mean, because its own precision is poor compared with the spread of true effects, so its posterior estimate lies close to 0.33. The large study's estimate stays near its own value, at about 0.26, since its precision is high relative to the spread. In general, the weight on a study's own data is the ratio of the between-study variance to the sum of the between-study and within-study variances. The figures are chosen to illustrate the logic and are not data from a real review.

Software

Bayesian meta-analysis can be carried out in R with packages built for the task, such as bayesmeta, and with general-purpose packages that interface to Stan, and in JAGS or Stan directly, which gives complete control over the model. Specialized packages for the normal-normal model are quick and need no simulation. General engines are needed for binomial likelihoods, network models and extensions. Whatever is used, the model code, the priors, the seeds and the software versions are recorded and shared so that the analysis can be reproduced.

When it is useful and when it is not

The method is most useful when studies are few, where a prior for the heterogeneity stabilizes the estimate and the uncertainty is reported honestly; when data are sparse, because exact likelihoods avoid continuity corrections; when external evidence should be incorporated; when the model is complex, as in network and multivariate meta-analysis; and when decision makers want probabilities. It is less useful when there are many studies, where the data swamp the prior and the results match a conventional analysis, and when the audience is unfamiliar with the approach and no benefit is gained. A conventional analysis with a Bayesian one alongside as a sensitivity analysis is a sound compromise in many reports. See random-effects meta-analysis.

Bayesian network meta-analysis

Bayesian methods are especially common in network meta-analysis, because the models are complex, many parameters share structure, and the output of interest, such as the probability that each treatment is best or the ranking probabilities, is a function of the joint posterior and is easily derived from simulated draws. The principles are the same as for pairwise analysis, with priors for the basic parameters, the heterogeneity and any inconsistency parameters, and the same attention to diagnostics and sensitivity. The method is described in the page on network meta-analysis.

Reporting

A Bayesian meta-analysis is reported against PRISMA 2020 with additional detail. The report gives the model and likelihood, every prior with its justification, the software and version, the number of chains, the length of warm-up and sampling, the convergence diagnostics, the posterior summaries and the results of the sensitivity analysis to the priors. The language should be understandable to readers without specialist training, so that the methods section explains what a credible interval is and how it differs from a confidence interval. The code and data are shared where possible, so that the analysis can be reproduced.

Limitations

The results depend on the priors, and an informative prior can be viewed as injecting opinion. The remedy is transparency and sensitivity analysis. With few studies, the posterior for heterogeneity largely reflects the prior, and no method can recover information that the data lack. Fitting is more demanding than for conventional analysis and calls for diagnostics. And, as with any meta-analysis, the quality of the included studies limits the conclusions.

How we can help

Support

Support for bayesian meta-analysis

The method can be supported at different depths. Choose what you need, and the scope is agreed in writing before work begins.

Get a quoteTell us your question, the number of studies and your target journal.

Frequently asked questions

What is a prior distribution?

A probability distribution that describes what is believed about an unknown quantity before seeing the data. It is combined with the data, through the likelihood, to give the posterior distribution.

How is a credible interval different from a confidence interval?

Given the model and priors, a 95 percent credible interval contains the true value with probability 95 percent. A confidence interval is defined through repeated sampling and does not have this direct probability interpretation.

Which prior should I use for heterogeneity?

A weakly informative prior on the standard deviation, such as a half-normal or half-Cauchy, is a common default, and empirical priors can be used when relevant data exist. Always report the prior and test alternatives.

What is MCMC and why must I check it?

Markov chain Monte Carlo simulates draws from the posterior. If the chains have not converged, the draws do not represent the posterior and the results can be wrong, so convergence diagnostics are required.

Do Bayesian and frequentist meta-analyses give different answers?

With many studies and weak priors they give very similar answers. They differ most with few studies, where priors matter, and in the way uncertainty is expressed.

Can I use informative priors from earlier studies?

Yes, where they are relevant and the choice is justified and transparent, with sensitivity analysis showing how much they affect the result. Double counting studies that are also in the data must be avoided.

References

  1. Sutton AJ, Abrams KR. Bayesian methods in meta-analysis and evidence synthesis. Stat Methods Med Res. 2001;10(4):277-303.
  2. Smith TC, Spiegelhalter DJ, Thomas A. Bayesian approaches to random-effects meta-analysis: a comparative study. Stat Med. 1995;14(24):2685-2699.
  3. Higgins JPT, Whitehead A. Borrowing strength from external trials in a meta-analysis. Stat Med. 1996;15(24):2733-2749.
  4. Spiegelhalter DJ, Abrams KR, Myles JP. Bayesian Approaches to Clinical Trials and Health-Care Evaluation. Wiley; 2004.
  5. Gelman A. Prior distributions for variance parameters in hierarchical models. Bayesian Anal. 2006;1(3):515-533.
  6. Gelman A, Carlin JB, Stern HS, Dunson DB, Vehtari A, Rubin DB. Bayesian Data Analysis. 3rd ed. CRC Press; 2013.
  7. Gelman A, Rubin DB. Inference from iterative simulation using multiple sequences. Stat Sci. 1992;7(4):457-472.
  8. Turner RM, Davey J, Clarke MJ, Thompson SG, Higgins JPT. Predicting the extent of heterogeneity in meta-analysis, using empirical data from the Cochrane Database of Systematic Reviews. Int J Epidemiol. 2012;41(3):818-827.
  9. Friede T, Rover C, Wandel S, Neuenschwander B. Meta-analysis of few small studies in orphan diseases. Res Synth Methods. 2017;8(1):79-91.
  10. Rover C. Bayesian random-effects meta-analysis using the bayesmeta R package. J Stat Softw. 2020;93(6):1-51.
  11. Carpenter B, Gelman A, Hoffman MD, et al. Stan: a probabilistic programming language. J Stat Softw. 2017;76(1):1-32.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.