Guide

Risk-of-bias tools

Risk-of-bias tools structure the judgment of whether a study's design, conduct or reporting is likely to have distorted its result. Which tool to use depends on the study design. This guide compares the main tools, explains how they are applied and how the results are presented and used.

What a risk-of-bias tool does

Studies are not equally trustworthy. Flaws in how participants were allocated, how outcomes were measured, how missing data were handled or how results were reported can push a result away from the truth. A risk-of-bias tool is a structured way of examining those features, one domain at a time, and reaching a judgment about whether each is likely to have biased the result. The judgment concerns internal validity: whether the study measured what it set out to measure. It is not a judgment of the study's importance, its size or whether it was well reported, although poor reporting often prevents a judgment.

The assessments serve three uses in a review. They are presented, so that readers can see the quality of the evidence. They feed sensitivity analyses, such as pooling only the studies at low risk. And they inform the rating of the certainty of evidence, where risk of bias is one of the reasons to lower confidence. The process is part of the systematic review method, and the service is described on the risk-of-bias and certainty assessment page.

Domains, not scores

Older approaches summarized quality as a single score, by adding points for features such as randomization, blinding and withdrawals. Such scales have fallen out of favor. They weight items arbitrarily, they mix features that bear on bias with others that do not, and a high total can hide a fatal flaw in one area. Studies have shown that conclusions can change depending on which scale is used. Current tools instead assess separate domains, each tied to a specific mechanism of bias, and reach a judgment for each, which are then combined into an overall judgment by explicit rules. The reader sees where the weaknesses are, and the review can respond to them. The domain-based approach is the recommended standard in the Cochrane Handbook and is used by most guideline groups.

The main tools

Risk-of-bias and critical appraisal tools by design
DesignToolDomains or content
Randomized trialsRoB 2Randomization process; deviations from intended interventions; missing outcome data; measurement of the outcome; selection of the reported result
Non-randomized studies of interventionsROBINS-IConfounding; selection of participants; classification of interventions; deviations from intended interventions; missing data; measurement of outcomes; selection of the reported result
Cohort and case-control studiesNewcastle-Ottawa ScaleSelection, comparability and outcome (or exposure), scored by stars
Diagnostic accuracy studiesQUADAS-2Patient selection; index test; reference standard; flow and timing
Prognostic factor studiesQUIPSParticipation; attrition; prognostic factor measurement; outcome measurement; confounding; statistical analysis and reporting
Prediction model studiesPROBASTParticipants; predictors; outcome; analysis
Systematic reviewsAMSTAR 2 or ROBISConduct and risk of bias of a review, as used in umbrella reviews
Other designsJBI checklists, CASP checklistsDesign-specific appraisal questions, including qualitative and cross-sectional studies

Tools are updated and new ones appear, so the review should name the version used and check the current guidance of the developers.

RoB 2 for randomized trials

RoB 2 is the revised Cochrane tool. It is applied to a specific result, not to a trial as a whole, because a trial can be at low risk for one outcome and high for another. For each of the five domains there is a set of signaling questions, answered yes, probably yes, probably no, no or no information, and an algorithm that maps the answers to a judgment of low risk, some concerns or high risk. The overall judgment for the result is the worst of the domain judgments, so a high risk in one domain makes the result high risk overall. The tool distinguishes the effect of assignment to the intervention, the intention-to-treat effect, from the effect of adhering to it, which affects the questions on deviations from intended interventions. Variants exist for cluster-randomized and crossover trials.

Applying the tool well requires the trial protocol, registry entry and statistical analysis plan in addition to the paper, since the fifth domain asks whether the reported result was selected from several possibilities, and that can only be judged against the plan. See the guide to GRADE for how the judgments feed into certainty.

ROBINS-I for non-randomized studies

ROBINS-I assesses a non-randomized study by comparing it with an ideal, hypothetical randomized trial of the same question, the target trial. Bias is judged as the departure from what that trial would have shown. It has seven domains, two before the intervention, one at the intervention, and four after it, and a judgment scale of low, moderate, serious and critical risk, with no information as an option. A study can be at low risk of bias only if it is comparable to a well-conducted randomized trial, which is rare for non-randomized studies, and confounding is the domain that most often leads to a serious rating. The assessor lists in advance the important confounders and co-interventions for the question, and checks whether the study measured and adjusted for them.

The tool is more demanding than RoB 2, requires content knowledge to specify the confounders, and takes longer. Many reviewers use simpler checklists for observational studies, such as the Newcastle-Ottawa Scale, which is quick but gives only a limited view of bias, is scored by summing stars, and has been criticized for inconsistent application. Reviews that rely on non-randomized evidence for decisions are better served by ROBINS-I.

Tools for other designs

QUADAS-2 is the standard for diagnostic accuracy studies, with four domains rated for risk of bias, three of which are also rated for concerns about applicability to the review question. QUIPS and the later PROBAST address the particular issues of prognosis: attrition, the measurement of the prognostic factor, confounding and the statistical analysis for factor studies, and the overfitting, handling of missing data and sample size in prediction models. AMSTAR 2 and ROBIS assess reviews, used when reviews are the unit of synthesis. Studies of prevalence, qualitative studies, cross-sectional studies and economic evaluations have their own appraisal tools, from sources such as the Joanna Briggs Institute and the CASP program. When no suitable tool exists, a review may adapt one, which should be documented and justified, and the adaptation reported. See diagnostic accuracy meta-analysis and prognostic meta-analysis.

Applying a tool in practice

  1. Choose the tool in the protocol

    Name the tool and the version for each design, and say how the judgments will be used, for example in sensitivity analyses.

  2. Calibrate

    The assessors apply the tool to a few studies together, discuss their reading of the signaling questions, and agree interpretations for the topic.

  3. Assess independently

    Two assessors judge each study and outcome independently, using all available sources, not only the published paper.

  4. Record support for each judgment

    The text of the study that led to the judgment is quoted, so that others can check it.

  5. Resolve differences

    Disagreements are settled by discussion or a third assessor, and the level of agreement may be reported.

  6. Use the results

    The judgments are presented, used in sensitivity analyses and carried into the certainty assessment.

Presenting the results

The standard displays are the traffic-light plot, which shows each study's judgment for each domain as a colored symbol, and the weighted summary bar, which shows the proportion of the evidence at each level of risk for each domain. A useful summary is the share of the analysis weight that comes from studies at each level. An invented example illustrates it. Six studies carry weights of 22, 18, 9, 31, 11, 9 percent, and are judged low, some concerns, high, low, high, some concerns. Studies at low risk of bias carry 53 percent of the weight, those with some concerns 27 percent and those at high risk 20 percent, although by number the split is 2 low, 2 some concerns and 2 high. The weighted view matters because a result dominated by a few biased studies is less trustworthy than a count of studies suggests. Tools for drawing these plots, such as the robvis package, are available. The judgments for every study and domain, with their supporting text, are provided as supplementary material.

Using the assessments in the synthesis

An assessment that does not affect the review is decoration. Three uses are standard. The first is a sensitivity analysis that excludes studies at high risk, or that restricts to those at low risk, to see whether the conclusion changes. The second is a subgroup or meta-regression analysis by risk of bias, to see whether effects are larger in studies at higher risk, which would be evidence of bias. The third is the certainty rating: when the studies that contribute most of the weight have serious limitations, the certainty of the evidence is lowered. The conclusions should be worded in line with these results. A review that reports risk of bias in a figure and then interprets the pooled effect without regard to it has not used the assessment. See sensitivity analysis.

Reporting the assessment

The methods section names the tool and its version for each design, states how many assessors worked on each study and how disagreements were resolved, and says how the judgments were used. The results present the judgment for each study and domain, ideally in a figure, with a summary across studies and the text of the support for each judgment available as a supplement. If a tool was adapted, the changes are described. If assessment was done by one reviewer with a second checking a sample, the limitation is stated. PRISMA 2020 asks for these items under the methods and results of risk-of-bias assessment, and many journals check for them. A report that gives the judgments but not the reasons leaves readers unable to evaluate them, and reviewers often ask for the supporting detail, so preparing it from the start saves a revision round later. See PRISMA 2020.

Common mistakes

  • Using a quality score and a cut-off to include or exclude studies.
  • Assessing the study and not the outcome, when risk differs between outcomes.
  • Using the wrong tool, for example RoB 2 on a non-randomized study.
  • Judging from the paper alone, without checking registry entries and protocols.
  • Treating missing information as low risk, instead of recording no information or some concerns.
  • Single-assessor judgments without checking, where independent assessment was planned.
  • Presenting the figure and ignoring it in the conclusions.

Support

Risk-of-bias assessment with the appropriate tool, visual summaries and the link to certainty ratings is part of the risk-of-bias and certainty assessment service.

Get a quoteTell us your included studies and designs.

Frequently asked questions

Which tool should I use for randomized trials?

RoB 2, the revised Cochrane tool, is the standard. It is applied to each result, with five domains and an algorithm from signaling questions to a judgment.

Can I use the Newcastle-Ottawa Scale for observational studies?

It is widely used and quick, but gives only a limited view of bias. ROBINS-I is more thorough for non-randomized studies of interventions, and the choice should be stated in the protocol.

Why should I avoid a single quality score?

Scores weight items arbitrarily and can hide a serious flaw in one domain behind a high total. Domain-based judgments show where the problems are.

Do I assess each study or each outcome?

Each result. A study can be at low risk for one outcome and high for another, so the assessment is made for the outcome used in the synthesis.

How many people should assess risk of bias?

Two, independently, with disagreements resolved by discussion or a third person, which is the usual standard.

What if the paper does not give enough information?

Check the registry entry, protocol and supplements, and ask the authors where practical. If the information is unavailable, record it as such and do not assume low risk.

References

  1. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  2. Sterne JA, Hernan MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
  3. Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
  4. Hayden JA, van der Windt DA, Cartwright JL, Cote P, Bombardier C. Assessing bias in studies of prognostic factors. Ann Intern Med. 2013;158(4):280-286.
  5. Wolff RF, Moons KGM, Riley RD, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58.
  6. Shea BJ, Reeves BC, Wells G, et al. AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ. 2017;358:j4008.
  7. Whiting P, Savovic J, Higgins JPT, et al. ROBIS: a new tool to assess risk of bias in systematic reviews was developed. J Clin Epidemiol. 2016;69:225-234.
  8. Juni P, Witschi A, Bloch R, Egger M. The hazards of scoring the quality of clinical trials for meta-analysis. JAMA. 1999;282(11):1054-1060.
  9. McGuinness LA, Higgins JPT. Risk-of-bias VISualization (robvis): an R package and Shiny web app for visualizing risk-of-bias assessments. Res Synth Methods. 2021;12(1):55-61.
  10. Wells GA, Shea B, O'Connell D, et al. The Newcastle-Ottawa Scale (NOS) for assessing the quality of nonrandomised studies in meta-analyses. Ottawa Hospital Research Institute.
  11. Page MJ, McKenzie JE, Bossuyt PM, et al. The PRISMA 2020 statement: an updated guideline for reporting systematic reviews. BMJ. 2021;372:n71. doi:10.1136/bmj.n71

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.