ISSN-registered · Peer-reviewed · Open Access
JournalsAboutContact
Scientific Writing

Risk of Bias Assessment: Choosing RoB 2, ROBINS-I or Newcastle-Ottawa

DE By Directive Editorial Team, Directive Publications ·4 Oct 2026 ·7 min read
Risk of Bias Assessment: Choosing RoB 2, ROBINS-I or Newcastle-Ottawa

A risk of bias assessment judges whether flaws in each included study could distort its results. Pick the tool by design: randomised trials get RoB 2, non-randomised intervention studies ROBINS-I, diagnostic accuracy studies QUADAS-3 and prediction models PROBAST+AI. Newcastle-Ottawa covers cohort and case-control studies but gives no official overall grade. Report judgements by domain, not as a summary score.

Bias here means systematic error, and Cochrane asks reviewers to read "risk of bias" as risk of material bias: bias likely to affect the conclusions drawn from a study. Imprecision is separate: a small trial can be at low risk of bias yet give a wide estimate. Formal appraisal is one feature that sets a systematic review apart from a mini or narrative review, where appraisal is informal.

Match each study design to a risk of bias assessment tool

Use one tool per design, so a review of trials and cohort studies needs two. Never borrow one tool's judgement labels for another.

Included designTool (version)DomainsJudgement levels
Individually randomised, parallel-group trialRoB 2 (2019)Randomisation process; deviations from intended interventions; missing outcome data; measurement of the outcome; selection of the reported resultLow risk; some concerns; high risk
Cluster-randomised or crossover trialRoB 2 variant for that design (2021)As RoB 2, plus a domain on when participants were identified and recruited (cluster); period and carryover effects are addressed (crossover)As RoB 2
Non-randomised study comparing interventionsROBINS-I (2016)Confounding; selection of participants; classification of interventions; deviations from intended interventions; missing data; measurement of the outcome; selection of the reported resultLow; moderate; serious; critical; or no information
Cohort (follow-up) study of an exposureROBINS-E (2024)Confounding; measurement of the exposure; selection of participants; post-exposure interventions; missing data; measurement of the outcome; selection of the reported result. Each also judged for direction of bias and threat to conclusionsLow risk; some concerns; high risk; very high risk
Cohort or case-control study (see limits below)Newcastle-Ottawa ScaleSelection (up to 4 stars); comparability (up to 2); exposure or outcome (up to 3)Stars per item; no official overall categories
Diagnostic accuracy studyQUADAS-3 (2026)Participants; index test; target condition; analysisLow; high; insufficient information
Prediction model development or validationPROBAST+AI (2025)Participants and data sources; predictors; outcome; analysisLow; high; unclear (quality concerns for development, risk of bias for evaluation)
Case series or prevalence studyJBI (formerly the Joanna Briggs Institute) checklist for that designThat checklist's itemsSet by that checklist

The developers of the Quality Assessment of Diagnostic Accuracy Studies (QUADAS) tool now recommend QUADAS-3, published in February 2026. PROBAST+AI updates the 2019 Prediction model Risk Of Bias ASsessment Tool (PROBAST). For the Risk Of Bias In Non-randomised Studies (ROBINS) tools, Version 2 of ROBINS-I is still a draft (its latest revision was posted in November 2025), so cite the 2016 tool. The Cochrane Handbook prefers RoB 2 for trials in new Cochrane reviews but accepts the original 2008/2011 tool "for the time being". If you use the older tool, say why.

Report Newcastle-Ottawa stars by category, not as a quality score

The Newcastle-Ottawa Scale is a star system for case-control and cohort studies, distributed by the Ottawa Hospital Research Institute. The scale and its coding manual give no total-score cut-offs, so "good", "fair" and "poor" labels are not part of it; if you use them, cite the source. The comparability item asks you to pick "the most important factor" to control for, so name your chosen factors in the protocol. Confounding and bias in observational studies explains how to choose them. The institute reports content validity and inter-rater reliability as established, and criterion validity and intra-rater reliability as still under examination.

Summary quality scores are discouraged. When Jüni and colleagues scored the same 17 heparin trials with 25 scales, the choice of scale changed the answer. For six scales, only the low-quality trials showed a clear benefit of low-molecular-weight over standard heparin; for seven, only the high-quality trials did; the other 12 gave similar results in both groups. The 1999 study concluded that relevant methodological aspects should be assessed individually. The PRISMA 2020 explanation and elaboration also prefers per-domain assessments, and the QUADAS-2 developers said their tool should not be used to generate a summary quality score. The same reasoning argues against a score as an inclusion cut-off: the Cochrane Handbook chapter on bias deals with high-risk studies in the analysis.

How domain judgements combine into an overall judgement

Signalling questions, answered yes, probably yes, probably no, no or no information, lead to a judgement for each domain, and the domain judgements set the overall one. In RoB 2 an algorithm proposes each judgement, and you may override it with stated reasons.

ToolLowest categoryMiddle categoryWorst category
RoB 2Low risk: every domain lowSome concerns: some concerns in at least one domain, none highHigh risk: any domain high, or some concerns in several domains that substantially lower confidence
ROBINS-ILow: every domain lowModerate: every domain low or moderateSerious: any domain serious, none critical. Critical: any domain critical
QUADAS-3Low: every domain lowInsufficient information (for reports too thin to judge, not a level of risk): any domain, none highHigh: any domain high

In RoB 2 and ROBINS-I, the overall judgement is at least as severe as the worst domain. Under ROBINS-I, several moderate domains may justify an overall serious, and several serious domains may justify critical.

All three tools judge a single result or estimate, not a whole study, so the same study can be at low risk for one result and high risk for another. Under ROBINS-I, a result at critical risk should not be included in any synthesis.

Fix the appraisal tool and its rules in the protocol before assessing

Write these decisions into the review protocol, so that no judgement is shaped by the results:

  • The tool and version for each design, and any adaptation.
  • For RoB 2, the effect of interest: assignment to the intervention (intention-to-treat) or adherence to it (per-protocol).
  • For ROBINS-I, the important confounding domains and co-interventions.
  • How the judgements will be used in the analysis.

Pilot the tool on three to six papers first, so assessors calibrate their answers. Two people should then assess each result independently, with a disagreement process set in advance, and record the quotation or source behind every judgement. Cochrane does not recommend kappa statistics; exploring why assessors disagreed matters more.

Carry risk of bias judgements into the analysis and GRADE certainty

Show each study's per-domain judgements in a table or traffic-light plot, with the full signalling-question answers in a supplementary file. Cochrane recommends forest plots that show each study's judgements beside its result.

Cochrane describes restricting the primary analysis to studies at low risk of bias, with sensitivity analyses, or stratifying analyses by risk of bias. It discourages a narrative discussion alone when studies differ in risk. A non-significant difference between high-risk and low-risk subgroups does not show an absence of bias, because such comparisons typically have low power.

Grading of Recommendations Assessment, Development and Evaluation (GRADE) then rates certainty in the body of evidence for each outcome as high, moderate, low or very low. Risk of bias is one of five domains that can lower it, beside inconsistency, indirectness, imprecision and publication bias (explained in reporting negative and null results). In Cochrane's GRADE chapter, trial evidence starts high and non-randomised evidence starts low; with ROBINS-I, non-randomised studies may start high but should generally lose two levels. A serious concern usually costs one level, a very serious one two.

A PRISMA 2020 reporting checklist for risk of bias and certainty

PRISMA 2020 has 27 items, plus a separate 12-item checklist for abstracts; the "Abstract" rows below use its numbers. These rows carry risk of bias and certainty, and items 15 and 22 were new in 2020. The wording is paraphrased from the PRISMA 2020 checklist and its explanation and elaboration.

SectionItemWhat to report
Abstract5Methods used to assess risk of bias
Methods11Tool and version; domains; rules for any overall judgement; adaptations; how many assessors, whether independent, and how disagreements were resolved; contact with investigators; any automation tools
Methods13e, 13fSubgroup, meta-regression or sensitivity analyses that use the judgements
Methods14Methods for assessing risk of bias due to missing results (reporting biases)
Methods15Certainty system and version; rules for reaching each level
Results18Each study's judgement per domain and overall, justified, for example with quotations
Results20a, 20dRisk of bias among the studies in each synthesis; every sensitivity analysis result
Results21Risk of bias due to missing results for each synthesis assessed
Results22Certainty for each outcome, with reasons for rating down or up, wherever results appear
Abstract; Discussion9; 23bLimitations of the evidence, such as risk of bias, inconsistency and imprecision

A Methods paragraph for item 11 can follow this pattern. Replace each bracketed slot and delete what does not apply.

Risk of bias was assessed for each result of [outcomes] using RoB 2 ([version date]) for randomised trials and ROBINS-I (2016) for non-randomised studies of interventions. The effect of interest was [assignment to / adherence to] the intervention. Two reviewers assessed each result independently after piloting the tools on [number] studies; disagreements were resolved by [method]. Study investigators were [contacted / not contacted] for missing information. Overall judgements followed each tool's rules, and every departure is justified in [supplementary file]. Certainty of evidence for each outcome was rated with GRADE.

At Directive Publications, the article types page describes a systematic review as using "explicit, reproducible methods for searching, selecting, and appraising the evidence" and requires a completed PRISMA checklist and flow diagram for one. Our author guidelines ask you to match the manuscript to the relevant EQUATOR checklist. No Directive policy page names a risk of bias tool or asks for GRADE, so the choice is yours, and PRISMA items 11, 15, 18 and 22 ask you to report it. If you spot an error in this guide, report the problem to us.

Frequently Asked Questions

Can the Newcastle-Ottawa Scale be used to appraise a randomised trial?
No. The Ottawa Hospital Research Institute distributes the Newcastle-Ottawa Scale for case-control and cohort studies only. For randomised trials, the Cochrane RoB 2 tool, with its variants for cluster-randomised and crossover trials, is built for the job. A cohort-study scale has no domain for problems in the randomisation process.
Is QUADAS-2 still acceptable for a diagnostic accuracy review?
QUADAS-3, published in February 2026, is now the version its developers recommend. It replaces the Flow and Timing domain with an Analysis domain and judges each estimate rather than each study. QUADAS-2 (2011) is still widely seen in published reviews, and the QUADAS developers note that a modified QUADAS-2 is recommended for Cochrane diagnostic accuracy reviews. Whichever you use, name the version in the Methods and explain why you chose it.
Is a risk of bias tool the same as a reporting guideline?
No. A reporting guideline such as CONSORT 2025 or STROBE tells authors what to report about their own study, while a risk of bias tool such as RoB 2 or ROBINS-I is how a reviewer judges whether a study's results could be biased. The two run in parallel: STARD 2015 pairs with QUADAS-3 for diagnostic accuracy, and TRIPOD+AI with PROBAST+AI for prediction models. AMSTAR 2 (A MeaSurement Tool to Assess systematic Reviews) is different again, because it appraises systematic reviews rather than primary studies.
What is the difference between risk of bias and certainty of evidence?
Risk of bias is judged for each study result, domain by domain, with a tool such as RoB 2 or ROBINS-I. Certainty of evidence is judged for the whole body of evidence on one outcome, with a system such as GRADE (Grading of Recommendations Assessment, Development and Evaluation), and is rated high, moderate, low or very low. GRADE can rate certainty down for five reasons: risk of bias, inconsistency, indirectness, imprecision and publication bias. A large effect, a dose-response gradient or plausible confounding that would shrink the observed effect can raise certainty, usually for non-randomised studies only.
DE
Directive Editorial Team
Directive Publications

The editorial team at Directive Publications — an international open-access publisher of peer-reviewed medical and scientific journals.

Publishing your research?

Directive Publications is an open-access publisher — every article peer-reviewed, Crossref-registered, and free to read under CC BY 4.0.

Submit a manuscript →Read our Scientific Writing policy →