Meta-analysis and evidence synthesis for genomics and bioinformatics

Genomics and bioinformatics study DNA, RNA and proteins at scale and develop the computational methods to analyze them. Meta-analysis is routine in genome-wide association studies and common for gene expression and biomarker work, with specialized methods for very large numbers of tests. Reviews of this field have to handle multiple testing, population structure, platform differences and overlapping samples.

Evidence synthesis in genomics and bioinformatics

Genomics studies the structure, function and variation of genomes. Bioinformatics supplies the computational tools to store, align, annotate and analyze the data. Meta-analysis is deeply embedded in this field because single studies are rarely large enough to detect the small effects of common genetic variants, and consortia routinely combine association results from dozens of cohorts. Reviews in the field also summarize candidate-gene studies, expression signatures, biomarker panels and benchmarks of software tools.

These tasks differ from clinical meta-analysis in scale and in how bias arises. A genome-wide study tests hundreds of thousands to millions of variants, so correction for multiple testing is essential. Ancestry differences between samples can produce false associations. Expression and sequencing data depend on platform, batch and processing. Our methods follow meta-analysis and systematic review practice, adapted to these features. This page builds on the general guidance for life sciences. It is about evidence synthesis and does not include wet-laboratory work or clinical interpretation.

Genome-wide association meta-analysis

Common methods for combining association results
MethodInputNotes
Inverse-variance fixed-effectEffect sizes (beta or log odds ratio) and standard errorsStandard when cohorts measure the same effect on the same scale; efficient and widely used
Sample-size weighted z-scoreZ-scores, direction of effect and sample sizeUseful when effect sizes are on different scales, as in different phenotype measurements
Random-effectsEffect sizes and standard errorsAllows true effects to differ between cohorts; special versions exist for genetics
Cochran's Q and I-squaredSame as aboveTest and quantify heterogeneity of a variant across cohorts
P-value combination (Fisher, Stouffer)P-valuesWeak when direction matters; ignores effect size

In genome-wide meta-analysis, consortia combine summary statistics across cohorts. Fixed-effect inverse-variance weighting is the default because effect sizes of a given variant are expected to be similar across cohorts of the same ancestry. Genome-wide significance is conventionally set at a p-value of 5 times ten to the minus eight, a threshold that reflects the number of independent common variants in the genome. Meta-analysis across ancestries needs methods that allow effects and allele frequencies to differ. A review should report the quality control applied to each cohort (call rate, Hardy-Weinberg equilibrium, imputation quality), the genomic inflation factor and the plots used to check for population stratification.

Population structure and confounding

Allele frequencies differ between populations, and so do many traits for environmental reasons. If a study mixes people of different ancestries, a variant may be associated with a trait only because both vary with ancestry. Cohorts address this with principal components, mixed models or matching, and the genomic inflation factor indicates how much inflation remains. A review of association studies should record how each cohort handled structure and whether results are for a single ancestry group. Most genomic data are from people of European ancestry, so findings and polygenic scores often transfer poorly to other groups, which a review should state clearly.

Candidate-gene studies and replication

Many older studies tested one or a few variants chosen on biological grounds, in samples of a few hundred. Large genome-wide studies later failed to replicate most of those associations, which is a well-known lesson on small studies, flexible analysis and publication bias. A synthesis of candidate-gene literature should be read in that light: pooled estimates from small studies may reflect bias, and the best evidence often comes from large consortia. Reviews of candidate-gene studies should test for Hardy-Weinberg equilibrium in controls, small-study effects and the first-study effect, where the earliest study reports a larger association than later ones. They should use reporting guidance for genetic association studies.

Gene expression and omics meta-analysis

Combining expression experiments poses problems of platform, normalization and batch. Approaches include combining effect sizes for each gene across studies, combining p-values or ranks, and merging normalized data after batch correction. Merged analyses can be affected by residual batch effects, which are particularly dangerous if batch is confounded with the condition of interest. A review should document the identification of datasets from repositories, the criteria for inclusion, the normalization method for each platform, quality checks and the strategy for multiple testing, with control of the false discovery rate. Pathway or gene set analyses have their own sources of bias, because annotated genes are those most studied, and the tests should account for gene length and other covariates where appropriate.

Single-cell and spatial data add scale and complexity, and meta-analysis methods are still developing. We describe their limits and avoid overconfident claims.

Biomarker discovery and validation

Studies of biomarker signatures often report impressive accuracy in discovery samples that fails in independent validation, because of overfitting, small samples and batch effects. A synthesis should separate discovery from validation, report validation cohorts with their accuracy and calibration, and use the methods for diagnostic accuracy and prognostic models described on the diagnostic accuracy and prognostic pages. Risk-of-bias tools for prediction models help to detect common flaws such as leakage of information between training and test sets.

Benchmarking and method comparison

Bioinformatics reviews often compare software tools for alignment, variant calling or differential expression. Benchmark studies differ in their data, in whether the truth is known (simulated or well-characterized reference samples), and in the metrics reported. Synthesis of benchmarks is hard because sets of tools and data differ, and results are sensitive to parameters. A review should describe benchmarks in a structured table, note whether the benchmark authors developed one of the tools compared and avoid ranking tools in general from results on specific data. Reproducibility depends on software versions and parameters, which should be recorded.

Sample overlap and shared controls

Consortia and biobanks reuse the same participants, and the same controls may appear in several studies. Overlap between samples causes correlation between results, which inflates the apparent evidence if ignored. Methods exist to estimate and correct for overlap using the correlation between test statistics. A review should check for shared cohorts and shared controls, particularly when combining published and consortium data, and exclude or adjust duplicate samples.

Polygenic scores and prediction

Polygenic scores sum the effects of many variants to predict a trait or disease risk. Studies report the variance explained, the area under the curve or the odds ratio between the top and bottom of the score distribution. These depend strongly on the ancestry of the discovery and target samples, on the trait and on the covariates used, and performance falls when scores are applied to groups unlike the discovery sample. A review should record the discovery and validation samples, whether they were independent, how variants were selected and weighted, and whether performance was reported with and without conventional risk factors. A statistically significant association is not the same as useful prediction, and the incremental value over simple clinical information is the more relevant question.

Reviews of this area should be cautious about causal claims. Methods that use genetic variants as instruments for exposures, called Mendelian randomization, rely on assumptions that can fail, and the review should note the checks done in the studies, such as tests for pleiotropy, and the number of variants used, since a result from a handful of variants is more fragile than one from hundreds.

Common pitfalls we look for

  • Using nominal p-values without correction for the number of tests.
  • Mixing ancestries without handling structure.
  • Treating small candidate-gene associations as established.
  • Merging expression data with batch confounded with condition.
  • Counting overlapping cohorts and shared controls as independent.
  • Reporting biomarker accuracy from discovery samples alone.

Planning a genomics synthesis

We help define the question (association, expression, biomarker or method comparison), the data sources (published studies, repositories such as GEO and the GWAS Catalog, consortium summary statistics), the inclusion criteria and the quality-control steps, plan the search and set up coding of ancestry, platform, sample size, covariates and multiple-testing control. Analysis is done in reproducible code. See the meta-analysis service for scope and process.

An invented example of fixed-effect combination

Suppose two invented cohorts estimate the effect of a variant on a trait: a beta of 0.10 with standard error 0.04 in the first, and 0.06 with standard error 0.03 in the second. The inverse-variance weights are 1 divided by 0.0016, which is 625, and 1 divided by 0.0009, which is about 1,111. The combined beta is (625 times 0.10 plus 1,111 times 0.06) divided by 1,736, which is about 0.0744, with standard error 1 divided by the square root of 1,736, about 0.024, and a z-statistic of about 3.1. That would be well short of the genome-wide threshold, which needs a z of about 5.45, a reminder that many variants cannot be detected without very large samples. Heterogeneity between the two is small, since the difference of 0.04 has a standard error of 0.05.

Data governance and ethics

Genomic data are sensitive. We work with published summary statistics and public repositories whose terms permit reuse and do not handle individual-level genotype data in this service. Results should not be used to identify individuals or to make clinical or ancestry claims about people.

How we support research projects in this area

Support

From experiments to a published synthesis

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    A structured question, effect-size choice, plan for dependent effects and registration where a platform accepts it.

  • Searching and extraction

    Searches across biological databases, extraction from text, tables and figures with checks between extractors.

  • Multilevel analysis

    Multilevel, phylogenetic and meta-regression models, with publication-bias analysis adapted to dependent data.

  • Manuscript and submission

    PRISMA-EcoEvo or PRISMA 2020 checklists, the manuscript, data and code for sharing.

Get a quoteDescribe your question, the kind of experiments and your target journal.

Boundaries of this service

A genomics and bioinformatics synthesis describes published associations, signatures and benchmarks. It does not provide clinical genetic interpretation, genetic counseling, variant classification or predictions for any person, and does not include laboratory work. Most data come from people of European ancestry, so findings may not transfer to other populations.

Frequently asked questions

How are GWAS results combined?

Usually by fixed-effect inverse-variance weighting of effect sizes and standard errors, or sample-size weighted z-scores, with quality control, a genomic inflation check and heterogeneity tests.

Why is the significance threshold so strict?

Because many variants are tested. The conventional genome-wide threshold, 5 times ten to the minus eight, reflects the number of independent common variants.

What went wrong with early candidate-gene studies?

Small samples, flexible analysis and publication bias produced associations that large studies mostly did not replicate.

Can expression datasets from different platforms be merged?

With care. Normalization and batch correction are essential, and batch confounded with condition can create false findings.

How do you handle overlapping cohorts?

By checking for shared samples and controls, excluding or adjusting duplicates and using methods that correct for correlation between statistics.

Do you interpret genetic results for individuals?

No. The service provides research and evidence-synthesis support only.

References

  1. Willer CJ, Li Y, Abecasis GR. METAL: fast and efficient meta-analysis of genomewide association scans. Bioinformatics. 2010;26(17):2190-2191.
  2. Hirschhorn JN, Lohmueller K, Byrne E, Hirschhorn K. A comprehensive review of genetic association studies. Genet Med. 2002;4(2):45-61.
  3. Ramasamy A, Mondry A, Holmes CC, Altman DG. Key issues in conducting a meta-analysis of gene expression microarray datasets. PLoS Med. 2008;5(9):e184.
  4. Martin AR, Kanai M, Kamatani Y, Okada Y, Neale BM, Daly MJ. Clinical use of current polygenic risk scores may exacerbate health disparities. Nat Genet. 2019;51(4):584-591.
  5. Little J, Higgins JPT, Ioannidis JPA, et al. Strengthening the reporting of genetic association studies (STREGA): an extension of the STROBE statement. PLoS Med. 2009;6(2):e22.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.