Guide

GRADE: rating the certainty of evidence

GRADE is a structured approach to rating how certain we can be about the effect of an intervention or exposure on an outcome, and, where guidelines are developed, the strength of recommendations. It is used by many guideline groups and systematic review authors. This guide explains the ratings, the domains and how results are presented.

What GRADE is

GRADE stands for Grading of Recommendations Assessment, Development and Evaluation. It was developed by a working group of methodologists and clinicians to replace a profusion of inconsistent systems for grading evidence, and it has been adopted by a large number of organizations and journals. It addresses two linked questions. How certain are we that the estimate of an effect reflects the true effect, which is the certainty of evidence? And how strongly should a recommendation be made, which depends on that certainty and on other factors, such as the balance of benefits and harms, values, resources and feasibility? This guide concerns the first, which systematic review authors usually rate, and touches on the second.

Two features set GRADE apart. The certainty is rated for each outcome and not for a study or a whole review, because evidence for one outcome can be strong while that for another is weak. And it is a structured judgment: the system names the factors to consider, requires reasons to be recorded, and presents the result transparently, but it does not turn the assessment into a calculation. Reasonable people can differ, and the system makes the difference visible.

The four levels

High
We are very confident that the true effect lies close to the estimate.
Moderate
We are moderately confident. The true effect is likely to be close to the estimate, but it may be substantially different.
Low
Our confidence is limited. The true effect may be substantially different from the estimate.
Very low
We have very little confidence. The true effect is likely to be substantially different from the estimate.

The ratings describe confidence in the estimate of effect, and not whether a treatment is good. Evidence of high certainty that an intervention has no effect is a finding of certainty. Evidence of low certainty that it works is a finding of uncertainty. The labels are used in summary tables and in the wording of conclusions.

Where the rating starts

In the standard approach, the starting point depends on the design of the studies. Evidence from randomized trials starts as high certainty, because randomization protects against many biases. Evidence from observational studies starts as low certainty, because confounding and selection bias are likely. The rating is then moved up or down by the considerations below. A body of randomized trials can fall to low or very low certainty if it has serious problems, and a body of observational studies can be raised to moderate or high if it has, for example, a very large effect. When a risk-of-bias tool for non-randomized studies such as ROBINS-I is used, some GRADE guidance starts the evidence at a higher level and then lowers it for risk of bias, to avoid counting the same limitation twice, and guidance on this point has been evolving.

Five reasons to lower certainty

Risk of bias
Limitations in the design or conduct of the studies that contribute most of the weight, such as inadequate concealment, lack of blinding for subjective outcomes, large losses to follow-up and selective outcome reporting. The tools described in the guide to risk-of-bias tools inform the judgment.
Inconsistency
Unexplained differences in the results of the studies, shown by wide variation in point estimates, little overlap of intervals and high heterogeneity statistics. Explained differences, for example by a subgroup, do not necessarily lower the rating. See heterogeneity.
Indirectness
Differences between the evidence and the question in the population, intervention, comparator or outcome, for instance surrogate outcomes, or reliance on indirect comparisons.
Imprecision
Wide confidence intervals that include both appreciable benefit and appreciable harm or no effect, or small numbers of participants or events. The rating considers the optimal information size, the sample that would be required for a single adequately powered trial.
Publication bias
A likelihood that studies are missing, suggested by funnel plot asymmetry, by a body of small positive studies, or by sponsorship. It is difficult to detect, and is often suspected and not proven. See publication bias.

Each factor can lower the rating by one level if serious, or by two levels if very serious. A body of evidence can be lowered by several factors, down to very low. The rater records the reason for each decision.

Three reasons to raise certainty

For observational evidence, three considerations can raise the rating. A large magnitude of effect, for example a relative risk above about 2 or below about 0.5 from studies with no major threats to validity, makes it unlikely that confounding alone could explain the result, and a very large effect, above about 5 or below about 0.2, can raise the rating by two levels. A dose-response gradient, in which the effect increases with the level of exposure, supports a causal interpretation. And when all plausible residual confounding would have reduced the observed effect, or would have produced an effect where none was observed, the rating can rise, since the true effect is probably at least as large as the estimate. These considerations are applied only to evidence that has no serious reason to be lowered, and they are applied with caution.

Summary of findings tables

The standard way of presenting GRADE results is the summary of findings table, which lists the important outcomes with, for each, the number of studies and participants, the relative effect, the absolute effects in the control and intervention groups, the certainty rating and comments or footnotes explaining the rating. Absolute effects are crucial, because relative effects can mislead, and they are calculated for stated baseline risks. An example shows the arithmetic. If the pooled risk ratio for an outcome is 0.80, with an interval from 0.65 to 0.98, and the baseline risk in the control group is 10 percent, then 100 per 1,000 people are expected to have the outcome without the intervention. With the intervention, the expected number is 80 per 1,000, which is 20 fewer, and the interval corresponds to between 2 and 35 fewer per 1,000. The number needed to treat is 1 divided by the absolute risk reduction of 0.02, or about 50. If the baseline risk were 1 percent, the same risk ratio would give only 2 fewer per 1,000, which shows why the baseline matters. The figures are invented for illustration.

The footnotes give the reasons for every downgrade or upgrade, which lets readers see and challenge the judgments. Software for building these tables is available, and some journals and Cochrane reviews require them.

A short illustration of downgrading

Imagine a body of evidence for an outcome from six randomized trials. Four of them did not conceal allocation and relied on a subjective outcome assessed by unblinded staff, and these four carry most of the weight. The reviewers judge risk of bias to be serious and lower the certainty by one level, from high to moderate. The point estimates range from a clear benefit to no effect, with intervals that barely overlap and no subgroup that explains the spread, so inconsistency is judged serious and the rating drops a second level, to low. The trials enrolled adults, and the question concerns children, so indirectness might lower it again, although here the reviewers judge that the mechanism is likely to apply and do not downgrade. The pooled interval includes both a meaningful benefit and no effect, but the total number of participants is above the optimal information size, so imprecision is not downgraded. Funnel plot asymmetry is absent, and there is no sign of missing studies. The final rating is low certainty, with footnotes giving the reason for each of the two downgrades. A reader can see the logic, disagree with a step, and see what would change the rating, for example a new, well-conducted trial. The scenario is invented to illustrate the process.

Choosing the outcomes to grade

Because ratings are made for outcomes, the choice of outcomes controls what a review can say. GRADE recommends that reviewers decide, before they look at the evidence, which outcomes are critical and important for the question, rating them in importance from the perspective of those affected. Outcomes that matter to patients, such as survival, function, symptoms, quality of life and serious harms, are given priority over surrogate or laboratory markers, unless the markers are known to predict the outcomes that matter. A review should include the adverse outcomes, not only the benefits, since a balanced view needs both. Limiting the summary of findings table to a handful of the most important outcomes keeps it usable, and the other outcomes are reported in the main text. Fixing the list in the protocol also guards against choosing outcomes after the results are known.

How a rating is made

  1. Choose the important outcomes

    The outcomes that matter most to patients and decision makers, including harms, are identified when the question is framed and before the evidence is examined.

  2. Start from the study design

    The body of evidence for each outcome is assigned its starting level.

  3. Consider each domain

    The reviewers assess risk of bias, inconsistency, indirectness, imprecision and publication bias, and, for observational evidence, the reasons to raise the rating.

  4. Decide and record

    A final rating is made, with a reason for every change from the starting point. Disagreements are discussed.

  5. Present

    The ratings are shown in the summary of findings table and reflected in the wording of the conclusions.

Wording conclusions

The certainty of evidence should shape how conclusions are written. Recommended phrasing uses standard verbs, such as improves or reduces, for high certainty evidence, probably improves for moderate, may improve for low, and uncertain whether it improves for very low certainty, with similar graded wording for no effect. This helps avoid overstating what the evidence shows. A conclusion that an intervention reduces mortality is a strong statement and should rest on strong evidence. Guidance from the GRADE working group on communicating the findings of systematic reviews describes the options. Writing conclusions in this way is as important as assigning the ratings, since readers take their impression from the words.

Extensions and related approaches

GRADE has been adapted to several contexts. For network meta-analysis, approaches such as CINeMA assess the contribution of each source of evidence to each comparison. For diagnostic tests, adaptations rate certainty for accuracy outcomes and for the effect of testing on patient outcomes. For qualitative evidence, GRADE-CERQual rates confidence in the findings of qualitative syntheses. For prognosis, guidance exists on rating certainty for prognostic factors and models. For guidelines, the Evidence to Decision frameworks bring together the certainty of evidence with other factors to reach recommendations. These adaptations share the structure of explicit domains and recorded reasons.

Criticisms and limitations

GRADE is widely used and also criticized. The judgments are subjective, and different groups can rate the same evidence differently. The categories are coarse, and a rating of low can cover a wide range. Starting observational evidence at low is regarded by some as too harsh and by others as appropriate. The guidance on imprecision and on inconsistency requires thresholds that are not universally agreed. The application is demanding in time, particularly when many outcomes are rated. And, since the system is organized around effects of interventions, it fits less well other kinds of question. Defenders reply that the explicitness of the process is its strength: it replaces opaque judgment with visible reasoning. The ratings describe confidence in an estimate of effect and are not recommendations for individual patients.

Certainty ratings are research outputs. They do not provide clinical advice.

Support

GRADE assessments and summary of findings tables are part of the risk-of-bias and certainty assessment service.

Get a quoteTell us your included studies and key outcomes.

Frequently asked questions

What does low certainty mean?

That our confidence in the estimate is limited and that the true effect may be substantially different from it. It is not a statement that the treatment does not work.

Can observational studies provide high-certainty evidence?

Yes, if they have no serious limitations and there is a reason to raise the rating, such as a very large effect or a dose-response gradient, though it is uncommon.

Is GRADE rated for each study?

No. It is rated for each outcome across the body of evidence.

Why do summary of findings tables show absolute effects?

Because relative effects can be misleading without knowing the baseline risk. Absolute effects show how many people are affected for a stated baseline.

Who does the rating?

Usually two reviewers independently, with disagreements resolved by discussion. The rating is a structured judgment and the reasons are recorded.

Is GRADE only for health?

It was developed for health. Its logic is used in some other fields, and adaptations exist for several types of evidence.

References

  1. Guyatt GH, Oxman AD, Vist GE, et al. GRADE: an emerging consensus on rating quality of evidence and strength of recommendations. BMJ. 2008;336(7650):924-926.
  2. Guyatt GH, Oxman AD, Schunemann HJ, Tugwell P, Knottnerus A. GRADE guidelines: a new series of articles in the Journal of Clinical Epidemiology. J Clin Epidemiol. 2011;64(4):380-382.
  3. Balshem H, Helfand M, Schunemann HJ, et al. GRADE guidelines: 3. Rating the quality of evidence. J Clin Epidemiol. 2011;64(4):401-406.
  4. Guyatt GH, Oxman AD, Kunz R, et al. GRADE guidelines: 8. Rating the quality of evidence: indirectness. J Clin Epidemiol. 2011;64(12):1303-1310.
  5. Guyatt GH, Oxman AD, Kunz R, et al. GRADE guidelines 6. Rating the quality of evidence: imprecision. J Clin Epidemiol. 2011;64(12):1283-1293.
  6. Guyatt GH, Oxman AD, Montori V, et al. GRADE guidelines: 5. Rating the quality of evidence: publication bias. J Clin Epidemiol. 2011;64(12):1277-1282.
  7. Guyatt GH, Oxman AD, Sultan S, et al. GRADE guidelines: 9. Rating up the quality of evidence. J Clin Epidemiol. 2011;64(12):1311-1316.
  8. Santesso N, Glenton C, Dahm P, et al. GRADE guidelines 26: informative statements to communicate the findings of systematic reviews of interventions. J Clin Epidemiol. 2020;119:126-135.
  9. Nikolakopoulou A, Higgins JPT, Papakonstantinou T, et al. CINeMA: an approach for assessing confidence in the results of a network meta-analysis. PLoS Med. 2020;17(4):e1003082.
  10. Lewin S, Booth A, Glenton C, et al. Applying GRADE-CERQual to qualitative evidence synthesis findings: introduction to the series. Implement Sci. 2018;13(Suppl 1):2.

Last updated October 2026. Methodological statements on this page follow the sources listed above.

Tell us about your research

Describe your question, study type and target journal. We will respond with the approach we would recommend and what we would need to begin.