Meta-analysis and evidence synthesis for political science

Political science studies governments, institutions, elections and political behavior. Its quantitative evidence includes field experiments on voting, survey experiments, cross-national panels and observational studies of institutions. Reviews have to handle cross-national data with few units, heterogeneous treatments, context effects and the sensitivity of findings to modelling choices.

Evidence synthesis in political science

Political science covers voting and public opinion, parties and campaigns, legislatures and executives, courts, democratization, conflict and political economy. Quantitative synthesis has taken hold where there are repeated tests of similar interventions, such as get-out-the-vote experiments, or where the same relationship has been estimated many times, such as the effect of electoral systems or of economic conditions on voting. Qualitative and historical approaches are central to the discipline, and a synthesis should respect them and not squeeze them into effect sizes.

Features of the evidence call for care. The units of analysis are often countries, parties or elections, so numbers are small and units are not interchangeable. Experiments in politics may not replicate across settings. Political topics provoke strong views, which makes transparency about inclusion and judgment especially important. Our methods follow systematic review and meta-analysis practice, adapted to these features. This page builds on the general guidance for social sciences. The service takes no political position.

Field and survey experiments

Types of experiment
TypeExampleIssue for synthesis
Field experimentMailers or door-to-door contact to raise voter turnoutMeasured behavior from official records; effects depend on baseline turnout, election type and message
Survey experimentFraming or information treatments in a questionnaireMeasures stated attitudes; short-lived effects; online samples
Lab experimentGames on trust or cooperationStudents; stakes small; transfer to politics unclear
Natural experimentLotteries, borders, close electionsLocal effects; assumptions about as-if randomness
Cross-national panelCountry-year data on institutions and outcomesFew countries; omitted variables; dependence over time

Get-out-the-vote research is a rare area where dozens of field experiments test similar tactics. Reviews have found that personal contact, such as face-to-face canvassing, raises turnout more than impersonal methods, with effects depending on the baseline turnout and the election. Treatment effects are reported in percentage points on turnout from official records. A review should extract the baseline rate, the mode of contact and the type of election, and examine whether effects are larger in low-turnout contexts. Survey experiments, by contrast, measure attitudes and often show effects that decay in days or weeks, so a review must distinguish attitude change from lasting persuasion.

Heterogeneity across contexts and treatments

Experiments of "the same" intervention vary in their message, messenger, country and political climate. Bayesian hierarchical models estimate how much true effects differ and allow predictions for new settings. A well-known project that ran coordinated experiments in several countries found that effects of an information treatment on voting behavior were small and varied across sites, a lesson in the limits of generalization. A review should present the between-study spread and use moderators on context cautiously, because treatment content is bound up with place.

Replication studies of prominent findings in political and social science have sometimes found smaller effects than the originals. A synthesis should look at whether early studies report larger effects than later or larger ones, and should include replications.

Cross-national data and institutions

Questions such as whether presidential systems are less stable, whether democracy promotes growth or whether electoral rules affect the number of parties are tested using panels of countries. Countries are few, relationships between variables are complicated, causation can run both ways, and results often depend on model specification, the sample of countries and the period. Synthesis in this area benefits from specification-curve and multiverse approaches, which show how estimates change under different reasonable modelling choices. A meta-analysis can code these choices and test whether they explain differences in results. Structural reasons for variation, such as differences in how democracy is measured by different indices, should be coded as well.

Measuring political concepts

Democracy, polarization, trust and ideology are measured by indices and survey scales that differ in content and method. Expert-coded indices differ from survey-based ones, and different indices of democracy can lead to different results. A review should record the measure and test whether findings depend on it. Survey items on the same attitude may be worded differently, in different languages and with different response scales, which limits comparability. Equivalence of measures across countries has to be assumed or tested, and reviews should state which.

Public opinion, media and misinformation

Research on persuasion, media effects, polarization and misinformation includes many survey experiments and some field experiments. Effects of single exposures on attitudes are typically small and short-lived, and effects of corrective information vary. Reviews should record the delay between treatment and outcome measurement, the nature of the sample, and whether the outcome is an attitude, a belief or a behavior. Online platforms present new data, such as observed browsing and sharing, with their own selection problems. The results of such studies depend on platform policies and times, and reviews should describe the period.

Conflict, security and institutions

Studies of civil conflict, peacekeeping, sanctions and terrorism use country-year data, event data and case studies. Event data are collected from media reports and so reflect media coverage as well as events. A review should describe the data source and its known biases. Peacekeeping studies, for instance, face the problem that missions are deployed to the hardest cases, so naive comparisons understate their effect. Designs that address selection are valuable, and a review should classify studies by whether they did so.

Publication bias and p-hacking

Political science has examined its own literature for evidence of publication bias, finding a visible jump in published p-values just below the conventional threshold in some areas. We apply funnel-based tests, p-curve and selection models, compare preregistered with other studies, and look at the role of replication projects. Registries exist for experiments in political science, and a review can compare registered with reported outcomes where records are available.

Courts, legislatures and administrative evidence

Studies of legislatures and courts use roll-call votes, judicial decisions and administrative records. Measures such as ideal points are estimated from votes with statistical models, so they are themselves uncertain, and treating them as fixed observations in second-stage analyses understates uncertainty. A review should note whether such uncertainty was propagated. Court decisions are a selected sample of disputes, since cases that settle never reach a judgment, and findings about judicial behavior apply to the decided cases. Comparative work on legislatures depends heavily on country context, and pooling across systems requires strong justification. Where a literature is mostly single-country, a systematic map is a better tool than a pooled estimate.

Administrative and bureaucratic performance studies use survey-based and record-based measures that correlate weakly. The review reports both and does not treat one as the truth.

Case studies and comparative historical evidence

Much political knowledge comes from case studies and historical comparison, which seek to explain outcomes through process tracing and comparison of a few cases. These methods cannot be pooled numerically, but a review can still map them, compare how studies define the concepts and conditions, and ask whether findings agree. Qualitative comparative analysis is sometimes used to examine combinations of conditions across cases. We describe such work in a structured narrative synthesis, using transparent criteria for including cases and for judging evidence, and linking it to the quantitative findings on the same question where possible. We do not present case study evidence as if it were a sample from which an average could be taken.

Common pitfalls we look for

  • Treating a few countries as a general sample.
  • Combining survey experiments and field experiments as if they measured the same outcome.
  • Ignoring baseline turnout or other baseline differences.
  • Using one index of democracy without checking others.
  • Treating event data as complete records.
  • Presenting a political conclusion as stronger than the evidence.

Planning a political science synthesis

We help define the question and units of analysis, decide whether pooling is appropriate, plan searches in Worldwide Political Science Abstracts, Scopus, Web of Science, SSRN, EGAP registries and working-paper series, and set up coding of design, context, measures, baseline rates and treatments. See the meta-analysis service for scope and process.

An invented example of a turnout effect

Suppose a review of door-to-door canvassing experiments finds an average increase in turnout of 3 percentage points across 40 invented experiments, with a prediction interval from 0 to 6. In an election where baseline turnout is 30 percent, 3 points would be a relative increase of 10 percent. For a mailer, the average might be 0.5 points. The difference between 3 and 0.5 points is large relative to either figure, and the cost per additional vote differs accordingly. The review would report effects by contact mode and baseline turnout, and would not claim that canvassing "works" in all settings, because studies mostly come from a few countries and election types.

Coding and transparency

Coding frames record country, election or period, unit of analysis, design, treatment content and mode, outcome and its source, baseline levels, sample, preregistration status and funding. Two coders extract data independently on a sample and agreement is reported. Because political topics are contested, we publish the protocol, inclusion decisions and analysis code so that readers with different views can check and challenge the work.

Neutrality and use of findings

Political findings are used selectively in public debate. A review states its question and rules in advance, discloses funders and takes no position on the policies or parties discussed. It reports uncertainty and negative or null findings with the same care as positive ones.

How we support research projects in this area

Support

From a policy or practice question to a published review

Support can cover a whole review or a single stage. The scope is agreed at the start.

  • Question and protocol

    Help choosing the review type, a protocol, and a plan for equity and context, with registration where suitable.

  • Searching and extraction

    Searches including grey literature, screening, and extraction for quantitative and qualitative strands.

  • Synthesis

    Meta-analysis, qualitative or mixed-methods synthesis, or an evidence map, as the question requires.

  • Manuscript and submission

    The report or manuscript, a policy summary, and journal or funder preparation.

Get a quoteDescribe the question, the designs you expect and the target journal or funder.

Boundaries of this service

A political science synthesis describes average findings across published studies. It does not forecast elections, advise campaigns, lobby or take political positions. Evidence comes from specific countries and periods, and many relationships are estimated from few units, so results may not generalize.

Frequently asked questions

Can field experiments on voting be meta-analyzed?

Yes. They report percentage point effects on turnout, and reviews examine variation by contact mode, baseline turnout and election type.

Why are cross-national studies hard to synthesize?

Few countries, omitted variables and sensitivity to model specification mean results often differ with reasonable modelling choices, which reviews code and test.

Do survey experiments show lasting persuasion?

Often not. Effects on stated attitudes can decay quickly, so reviews separate attitude change from durable shifts.

How do you deal with different indices of democracy?

By recording the index used and testing whether findings change under others.

How is neutrality maintained?

By a pre-specified protocol, transparent inclusion rules, disclosure of funding and equal treatment of null and positive results.

Do you provide election forecasts or campaign advice?

No. The service provides research and evidence-synthesis support only.

References

  1. Green DP, Gerber AS. Get out the vote: how to increase voter turnout. 4th ed. Washington, DC: Brookings Institution Press; 2019.
  2. Dunning T, Grossman G, Humphreys M, et al. Voter information campaigns and political accountability: cumulative findings from a preregistered meta-analysis of coordinated trials. Sci Adv. 2019;5(7):eaaw2612.
  3. Gerber AS, Malhotra N. Do statistical reporting standards affect what is published? Publication bias in two leading political science journals. Q J Polit Sci. 2008;3(3):313-326.
  4. Vivalt E. How much can we generalize from impact evaluations? J Eur Econ Assoc. 2020;18(6):3045-3089.
  5. Steegen S, Tuerlinckx F, Gelman A, Vanpaemel W. Increasing transparency through a multiverse analysis. Perspect Psychol Sci. 2016;11(5):702-712.
  6. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.