A risk of bias assessment judges whether flaws in each included study could distort its results. Pick the tool by design: randomised trials get RoB 2, non-randomised intervention studies ROBINS-I, diagnostic accuracy studies QUADAS-3 and prediction models PROBAST+AI. Newcastle-Ottawa covers cohort and case-control studies but gives no official overall grade. Report judgements by domain, not as a summary score.
Bias here means systematic error, and Cochrane asks reviewers to read "risk of bias" as risk of material bias: bias likely to affect the conclusions drawn from a study. Imprecision is separate: a small trial can be at low risk of bias yet give a wide estimate. Formal appraisal is one feature that sets a systematic review apart from a mini or narrative review, where appraisal is informal.
Match each study design to a risk of bias assessment tool
Use one tool per design, so a review of trials and cohort studies needs two. Never borrow one tool's judgement labels for another.
| Included design | Tool (version) | Domains | Judgement levels |
|---|---|---|---|
| Individually randomised, parallel-group trial | RoB 2 (2019) | Randomisation process; deviations from intended interventions; missing outcome data; measurement of the outcome; selection of the reported result | Low risk; some concerns; high risk |
| Cluster-randomised or crossover trial | RoB 2 variant for that design (2021) | As RoB 2, plus a domain on when participants were identified and recruited (cluster); period and carryover effects are addressed (crossover) | As RoB 2 |
| Non-randomised study comparing interventions | ROBINS-I (2016) | Confounding; selection of participants; classification of interventions; deviations from intended interventions; missing data; measurement of the outcome; selection of the reported result | Low; moderate; serious; critical; or no information |
| Cohort (follow-up) study of an exposure | ROBINS-E (2024) | Confounding; measurement of the exposure; selection of participants; post-exposure interventions; missing data; measurement of the outcome; selection of the reported result. Each also judged for direction of bias and threat to conclusions | Low risk; some concerns; high risk; very high risk |
| Cohort or case-control study (see limits below) | Newcastle-Ottawa Scale | Selection (up to 4 stars); comparability (up to 2); exposure or outcome (up to 3) | Stars per item; no official overall categories |
| Diagnostic accuracy study | QUADAS-3 (2026) | Participants; index test; target condition; analysis | Low; high; insufficient information |
| Prediction model development or validation | PROBAST+AI (2025) | Participants and data sources; predictors; outcome; analysis | Low; high; unclear (quality concerns for development, risk of bias for evaluation) |
| Case series or prevalence study | JBI (formerly the Joanna Briggs Institute) checklist for that design | That checklist's items | Set by that checklist |
The developers of the Quality Assessment of Diagnostic Accuracy Studies (QUADAS) tool now recommend QUADAS-3, published in February 2026. PROBAST+AI updates the 2019 Prediction model Risk Of Bias ASsessment Tool (PROBAST). For the Risk Of Bias In Non-randomised Studies (ROBINS) tools, Version 2 of ROBINS-I is still a draft (its latest revision was posted in November 2025), so cite the 2016 tool. The Cochrane Handbook prefers RoB 2 for trials in new Cochrane reviews but accepts the original 2008/2011 tool "for the time being". If you use the older tool, say why.
Report Newcastle-Ottawa stars by category, not as a quality score
The Newcastle-Ottawa Scale is a star system for case-control and cohort studies, distributed by the Ottawa Hospital Research Institute. The scale and its coding manual give no total-score cut-offs, so "good", "fair" and "poor" labels are not part of it; if you use them, cite the source. The comparability item asks you to pick "the most important factor" to control for, so name your chosen factors in the protocol. Confounding and bias in observational studies explains how to choose them. The institute reports content validity and inter-rater reliability as established, and criterion validity and intra-rater reliability as still under examination.
Summary quality scores are discouraged. When Jüni and colleagues scored the same 17 heparin trials with 25 scales, the choice of scale changed the answer. For six scales, only the low-quality trials showed a clear benefit of low-molecular-weight over standard heparin; for seven, only the high-quality trials did; the other 12 gave similar results in both groups. The 1999 study concluded that relevant methodological aspects should be assessed individually. The PRISMA 2020 explanation and elaboration also prefers per-domain assessments, and the QUADAS-2 developers said their tool should not be used to generate a summary quality score. The same reasoning argues against a score as an inclusion cut-off: the Cochrane Handbook chapter on bias deals with high-risk studies in the analysis.
How domain judgements combine into an overall judgement
Signalling questions, answered yes, probably yes, probably no, no or no information, lead to a judgement for each domain, and the domain judgements set the overall one. In RoB 2 an algorithm proposes each judgement, and you may override it with stated reasons.
| Tool | Lowest category | Middle category | Worst category |
|---|---|---|---|
| RoB 2 | Low risk: every domain low | Some concerns: some concerns in at least one domain, none high | High risk: any domain high, or some concerns in several domains that substantially lower confidence |
| ROBINS-I | Low: every domain low | Moderate: every domain low or moderate | Serious: any domain serious, none critical. Critical: any domain critical |
| QUADAS-3 | Low: every domain low | Insufficient information (for reports too thin to judge, not a level of risk): any domain, none high | High: any domain high |
In RoB 2 and ROBINS-I, the overall judgement is at least as severe as the worst domain. Under ROBINS-I, several moderate domains may justify an overall serious, and several serious domains may justify critical.
All three tools judge a single result or estimate, not a whole study, so the same study can be at low risk for one result and high risk for another. Under ROBINS-I, a result at critical risk should not be included in any synthesis.
Fix the appraisal tool and its rules in the protocol before assessing
Write these decisions into the review protocol, so that no judgement is shaped by the results:
- The tool and version for each design, and any adaptation.
- For RoB 2, the effect of interest: assignment to the intervention (intention-to-treat) or adherence to it (per-protocol).
- For ROBINS-I, the important confounding domains and co-interventions.
- How the judgements will be used in the analysis.
Pilot the tool on three to six papers first, so assessors calibrate their answers. Two people should then assess each result independently, with a disagreement process set in advance, and record the quotation or source behind every judgement. Cochrane does not recommend kappa statistics; exploring why assessors disagreed matters more.
Carry risk of bias judgements into the analysis and GRADE certainty
Show each study's per-domain judgements in a table or traffic-light plot, with the full signalling-question answers in a supplementary file. Cochrane recommends forest plots that show each study's judgements beside its result.
Cochrane describes restricting the primary analysis to studies at low risk of bias, with sensitivity analyses, or stratifying analyses by risk of bias. It discourages a narrative discussion alone when studies differ in risk. A non-significant difference between high-risk and low-risk subgroups does not show an absence of bias, because such comparisons typically have low power.
Grading of Recommendations Assessment, Development and Evaluation (GRADE) then rates certainty in the body of evidence for each outcome as high, moderate, low or very low. Risk of bias is one of five domains that can lower it, beside inconsistency, indirectness, imprecision and publication bias (explained in reporting negative and null results). In Cochrane's GRADE chapter, trial evidence starts high and non-randomised evidence starts low; with ROBINS-I, non-randomised studies may start high but should generally lose two levels. A serious concern usually costs one level, a very serious one two.
A PRISMA 2020 reporting checklist for risk of bias and certainty
PRISMA 2020 has 27 items, plus a separate 12-item checklist for abstracts; the "Abstract" rows below use its numbers. These rows carry risk of bias and certainty, and items 15 and 22 were new in 2020. The wording is paraphrased from the PRISMA 2020 checklist and its explanation and elaboration.
| Section | Item | What to report |
|---|---|---|
| Abstract | 5 | Methods used to assess risk of bias |
| Methods | 11 | Tool and version; domains; rules for any overall judgement; adaptations; how many assessors, whether independent, and how disagreements were resolved; contact with investigators; any automation tools |
| Methods | 13e, 13f | Subgroup, meta-regression or sensitivity analyses that use the judgements |
| Methods | 14 | Methods for assessing risk of bias due to missing results (reporting biases) |
| Methods | 15 | Certainty system and version; rules for reaching each level |
| Results | 18 | Each study's judgement per domain and overall, justified, for example with quotations |
| Results | 20a, 20d | Risk of bias among the studies in each synthesis; every sensitivity analysis result |
| Results | 21 | Risk of bias due to missing results for each synthesis assessed |
| Results | 22 | Certainty for each outcome, with reasons for rating down or up, wherever results appear |
| Abstract; Discussion | 9; 23b | Limitations of the evidence, such as risk of bias, inconsistency and imprecision |
A Methods paragraph for item 11 can follow this pattern. Replace each bracketed slot and delete what does not apply.
Risk of bias was assessed for each result of [outcomes] using RoB 2 ([version date]) for randomised trials and ROBINS-I (2016) for non-randomised studies of interventions. The effect of interest was [assignment to / adherence to] the intervention. Two reviewers assessed each result independently after piloting the tools on [number] studies; disagreements were resolved by [method]. Study investigators were [contacted / not contacted] for missing information. Overall judgements followed each tool's rules, and every departure is justified in [supplementary file]. Certainty of evidence for each outcome was rated with GRADE.
At Directive Publications, the article types page describes a systematic review as using "explicit, reproducible methods for searching, selecting, and appraising the evidence" and requires a completed PRISMA checklist and flow diagram for one. Our author guidelines ask you to match the manuscript to the relevant EQUATOR checklist. No Directive policy page names a risk of bias tool or asks for GRADE, so the choice is yours, and PRISMA items 11, 15, 18 and 22 ask you to report it. If you spot an error in this guide, report the problem to us.