Risk-of-bias and certainty assessment

We assess the risk of bias of included studies with the tool that fits each design, and rate the certainty of the evidence for each outcome with GRADE. These judgments determine how far a review's conclusions can be trusted, and they are among the parts of a review that editors and reviewers examine most closely.

Why bias and certainty are assessed

Studies are not equally trustworthy, and a review that treats them as if they were can reach a precise but wrong conclusion. Risk of bias is the possibility that the design, conduct or reporting of a study systematically distorts its result. Assessing it for each included study allows the review to take account of the quality of the evidence in the synthesis and in the conclusions, for example through sensitivity analyses that exclude studies at high risk. It is a judgment about the study's internal validity, not a score of its overall quality or its importance.

The certainty of the evidence is a related but wider judgment, made for each important outcome across the whole body of studies. It asks how confident we can be that the true effect is close to the estimate. Risk of bias is one input. Others include inconsistency between studies, whether the evidence directly addresses the question, the precision of the estimate and the likelihood of publication bias. The systematic review service places these assessments within the project.

What the service includes

  • Choice of tool for each study design, justified in the protocol.
  • Independent assessment of each included study by two reviewers, with disagreements resolved by discussion or a third reviewer.
  • Support for each judgment, quoting the text of the study that led to it, so that others can check it.
  • Visual summaries, such as traffic-light plots and weighted summary charts.
  • Use of the judgments in the synthesis, including sensitivity analyses.
  • GRADE assessment for each important outcome, with the reasoning for every rating.
  • Summary of findings table and evidence profile for the report.
  • Methods text describing the approach, ready for the manuscript.

Choosing the right tool

Risk-of-bias tools are designed for particular designs, and using the wrong one gives meaningless results. The main tools are listed here.

Risk-of-bias tools by study design
DesignToolWhat it examines
Randomized trialsRoB 2Randomization process, deviations from intended interventions, missing outcome data, outcome measurement, selection of the reported result
Non-randomized studies of interventionsROBINS-IConfounding, participant selection, classification of interventions, deviations, missing data, outcome measurement, selection of the reported result
Diagnostic accuracy studiesQUADAS-2Patient selection, index test, reference standard, flow and timing
Prognostic factor studiesQUIPSParticipation, attrition, prognostic factor and outcome measurement, confounding, statistical analysis
Prediction model studiesPROBASTParticipants, predictors, outcome and analysis
Systematic reviewsAMSTAR 2 or ROBISThe conduct and risk of bias of a review itself, for umbrella reviews

Older scales that summarize quality in a single score, and checklists such as the Newcastle-Ottawa Scale for observational studies, are still used and accepted by some journals, but domain-based tools are recommended because they separate the sources of bias and do not hide serious problems behind a high total. The guide to risk-of-bias tools compares them.

How assessments are made

The domain-based tools work through signaling questions about the features of a study that matter for each source of bias, such as whether the allocation sequence was concealed or whether outcome assessors knew the treatment. The answers lead to a judgment of low risk, some concerns or high risk for each domain and an overall judgment for the result. The judgments depend on what the study reports, and many studies report too little to judge. Where a published paper is unclear, the protocol, the trial registration, supplements and, where practical, correspondence with the authors are used.

Two assessors working independently is the norm, because judgments involve interpretation. The process also benefits from a calibration exercise on a few studies to align interpretation of the signaling questions. For each judgment, the supporting text from the study is recorded. The assessment concerns the specific result being used in the review and not the study as a whole, because a study can be at low risk for one outcome and at high risk for another.

Using the assessments in the review

The judgments should affect the review, otherwise they are decoration. They are presented in tables and figures that readers can examine. They feed sensitivity analyses that exclude studies at high risk of bias to see whether the conclusions change, and where there are enough studies, subgroup analyses by risk of bias. They contribute to the certainty ratings. And they shape the wording of the conclusions: a pooled estimate drawn largely from studies at high risk of bias is described with the appropriate caution. Ignoring bias in the interpretation while reporting it in a table is a common failing in published reviews.

GRADE: rating the certainty of evidence

GRADE is a structured approach to rating the certainty of the evidence for each outcome and, where recommendations are being made, the strength of recommendations. The certainty of evidence is rated as high, moderate, low or very low. Evidence from randomized trials begins at high certainty and from observational studies at low, and the rating is then lowered or raised according to defined domains. Five factors can lower it: risk of bias, inconsistency of results, indirectness of the evidence to the question, imprecision of the estimate and likelihood of publication bias. Three can raise the certainty of observational evidence: a large effect, a dose-response gradient, and the situation in which all plausible confounding would reduce the observed effect.

Each judgment is made for an outcome, because the certainty differs between outcomes in the same review, and each is accompanied by a reason. The ratings are presented in a summary of findings table, which shows for each important outcome the number of studies and participants, the relative and absolute effects and the certainty rating with footnotes explaining the decisions. The format makes the reasoning visible and gives readers a compact view of what the evidence shows and how much confidence it deserves. See GRADE.

The five reasons to lower certainty

Risk of bias
Serious limitations in the design or conduct of the studies that contribute most to the estimate. Concerns are judged across studies, weighted by their contribution.
Inconsistency
Unexplained differences in results between studies, judged from the spread of estimates, the overlap of their intervals and the heterogeneity statistics, in light of any explanation found.
Indirectness
Differences between the evidence and the question in population, intervention, comparator or outcome, or evidence from indirect comparisons.
Imprecision
Wide intervals that include both important benefit and important harm or no effect, or too few participants or events to be confident.
Publication bias
A likelihood that studies are missing, suggested by a funnel plot asymmetry, a pattern of small positive trials, or industry sponsorship, and interpreted cautiously because asymmetry has other causes.

Missing information and contacting authors

Many studies do not report enough to answer the signaling questions, for example how the allocation sequence was generated or whether the analysis followed the original plan. Before a judgment of "no information" is accepted, the study's registry entry, protocol, statistical analysis plan and supplements are checked, since they often contain what the paper omits. Where it is practical and proportionate, the authors are asked, with specific questions and a clear statement of what the information is for. The response, or its absence, is recorded. A study's lack of reporting is not the same as poor conduct, and the assessment should say which it is judging.

Bias across the body of evidence

Bias can arise at the level of the whole literature as well as within studies. A study may measure several outcomes and report only some, or analyze its data in several ways and report the most favorable, which is selective reporting of the result. Comparing the outcomes in a publication with those in the protocol or registry is the usual check. Across studies, the missing results of unpublished or selectively published work can distort a synthesis, which is the problem of publication bias. Funnel plots and related tests give a signal when enough studies are available, but asymmetry has other causes, so the interpretation is cautious. In GRADE, concern about these meta-biases is one of the reasons to lower the certainty of evidence. The guides on publication bias and small-study effects explain the issues.

Reporting

The methods text names the tools and versions used, the number of assessors, how disagreements were resolved, and how the judgments were used. The results present the judgments for each study and domain, ideally in a figure, with the summary of findings table for the main outcomes. PRISMA 2020 asks for the methods and results of risk-of-bias assessment and certainty assessment, and many journals expect them. Where the assessment was done on a sample or by one reviewer with a second checking, the limitation is stated. The full set of judgments with their supporting text is supplied as supplementary material.

Deliverables

  • Risk-of-bias judgments for each study and outcome, with supporting text.
  • Figures summarizing the judgments in publication-ready formats.
  • GRADE assessment with the reasoning for each domain and outcome.
  • Summary of findings table and evidence profile.
  • Sensitivity analyses based on risk of bias, if the synthesis is in scope.
  • Methods and results text for the manuscript.

Get a quoteTell us your included studies and key outcomes.

Optional extension: from assessment to published paper

Optional extension

Carry the assessment through to a full manuscript

When the scope continues beyond the assessment, the package also contains the following.

Get a quoteTell us your scope, target journal and deadline.

Manuscript preparation follows the research integrity and authorship statement. The researchers who conceived the study and interpret its findings remain responsible for the content and its conclusions, and contributions that do not meet authorship criteria are acknowledged.

Limitations

Risk-of-bias assessment depends on what studies report, and poor reporting is often mistaken for poor conduct. Judgments involve interpretation and are made differently by different assessors, which is why independent assessment and explicit reasoning matter. The tools identify the potential for bias and do not measure its size, so a rating of high risk does not mean the result is wrong, nor does a low rating mean it is right. GRADE ratings are judgments and not calculations, and reasonable assessors can differ, especially on inconsistency and imprecision. The system structures judgment and makes it transparent, and it does not remove it. Certainty ratings describe confidence in an estimate and are not recommendations for action.

This service provides research and evidence-synthesis support. It does not provide clinical advice, and certainty ratings are not recommendations for the care of individual patients.

How long does it take?

The time depends on the number of studies and outcomes, the design of the studies, how fully they are reported, and whether authors must be contacted. Two independent assessments of every study take longer than one, and GRADE ratings take time for each outcome. After the included studies and outcomes are known, a schedule and its dependencies are agreed.

Frequently asked questions

What is the difference between risk of bias and certainty of evidence?

Risk of bias concerns the validity of individual studies. Certainty of evidence is a judgment about the body of evidence for an outcome, to which risk of bias is one contribution alongside inconsistency, indirectness, imprecision and publication bias.

Why not use a single quality score?

Summary scores can hide serious problems in one area behind a high total, and they weight items arbitrarily. Domain-based tools show where bias arises and allow the review to respond to it.

How many people assess risk of bias?

Two working independently is the usual standard, with disagreements resolved by discussion or a third person. Documented alternatives are sometimes used, with their limits stated.

Can observational studies reach high certainty in GRADE?

They start lower than randomized trials, but certainty can be raised for a large effect, a dose-response gradient, or when plausible confounding would reduce the effect. Some approaches that use ROBINS-I begin from a higher level and then lower it.

Does a high-certainty rating mean I should recommend a treatment?

No. It describes confidence in the estimate of effect. Recommendations also weigh the balance of benefits and harms, values, costs and feasibility.

Can you assess studies in my own review?

Yes. The assessment can be done on the studies you have included, using the protocol and the tools it specifies.

References

  1. Sterne JAC, Savovic J, Page MJ, et al. RoB 2: a revised tool for assessing risk of bias in randomised trials. BMJ. 2019;366:l4898.
  2. Sterne JA, Hernan MA, Reeves BC, et al. ROBINS-I: a tool for assessing risk of bias in non-randomised studies of interventions. BMJ. 2016;355:i4919.
  3. Whiting PF, Rutjes AWS, Westwood ME, et al. QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Ann Intern Med. 2011;155(8):529-536.
  4. Hayden JA, van der Windt DA, Cartwright JL, Cote P, Bombardier C. Assessing bias in studies of prognostic factors. Ann Intern Med. 2013;158(4):280-286.
  5. Wolff RF, Moons KGM, Riley RD, et al. PROBAST: a tool to assess the risk of bias and applicability of prediction model studies. Ann Intern Med. 2019;170(1):51-58.
  6. Shea BJ, Reeves BC, Wells G, et al. AMSTAR 2: a critical appraisal tool for systematic reviews that include randomised or non-randomised studies of healthcare interventions, or both. BMJ. 2017;358:j4008.
  7. Whiting P, Savovic J, Higgins JPT, et al. ROBIS: a new tool to assess risk of bias in systematic reviews was developed. J Clin Epidemiol. 2016;69:225-234.
  8. McGuinness LA, Higgins JPT. Risk-of-bias VISualization (robvis): an R package and Shiny web app for visualizing risk-of-bias assessments. Res Synth Methods. 2021;12(1):55-61.
  9. Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924-926.
  10. Guyatt GH, Oxman AD, Kunz R, et al. GRADE guidelines: 2. Framing the question and deciding on important outcomes. J Clin Epidemiol. 2011;64(4):395-400.
  11. Balshem H, Helfand M, Schunemann HJ, et al. GRADE guidelines: 3. Rating the quality of evidence. J Clin Epidemiol. 2011;64(4):401-406.
  12. Santesso N, Glenton C, Dahm P, et al. GRADE guidelines 26: informative statements to communicate the findings of systematic reviews of interventions. J Clin Epidemiol. 2020;119:126-135.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.