Assessing bias in clinical trials is rarely as clean as the flowcharts suggest
I spend most of my time going through systematic reviews and watching people fill out bias assessment forms like they're checking boxes on a tax return. The process looks mechanical, but anyone who has actually sat with a trial protocol for more than an hour knows that decisions often come down to judgment calls you won't find spelled out in a manual. The Cochrane Risk Of Bias Assessment Tool version 2, often called RoB 2, is the current standard for randomized trials. It moved away from the old single-score system and now evaluates five domains: randomization process, deviations from intended interventions, missing outcome data, measurement of the outcome, and selective reporting. Each domain gets rated as low risk, some concerns, or high risk. That sounds straightforward until you open a real paper and realize the authors rarely describe their randomization sequence in enough detail to make any of those calls confidently.
Working through the Risk Of Bias Assessment Tool in practice
Start by pulling the published report and the trial registry entry side by side. A lot of people skip the registry part. I don't. The protocol registered on ClinicalTrials.gov or the ISRCTN portal often contains outcome definitions and analysis plans that the final publication quietly drops or reorders. The selective reporting domain depends entirely on noticing those mismatches. For the randomization domain, look for evidence that the sequence was concealed before allocation. Words like "computer-generated" or "random permuted blocks" tell you the sequence was unpredictable, but that is not the same as concealment. If the envelope system is not described, if dates of envelope preparation are missing, or if the method allowed someone to peek ahead, you flag it. I once spent three hours cross-referencing a cardiology trial only to discover that the supposed sealed envelope method was never independently audited and the clinic staff could see the allocation list on a shared whiteboard in the hallway. The trial was rated high risk for selection bias and I flagged it in the review within about twenty minutes of realizing what was going on. For intervention deviations, you need to account for blinding. If participants know they are in the placebo group, they may drop out or seek alternative treatments. If clinicians know the assignment, they may alter co-interventions. The tool asks you to consider whether these deviations created meaningful differences between groups. In a physical therapy trial I reviewed last year, the intervention group received six supervised sessions while the control group was told to exercise at home with no check-ins. The treatment fidelity gap was enormous, but the authors presented it as a straightforward comparison. The some concerns rating felt like the only honest choice.
Missing data is where most reviews fall apart. People look at the dropout percentage and move on. You should instead check whether the missingness is related to the outcome. If patients with severe adverse effects dropped out of the drug arm and the analysis only used completers, the result is almost certainly biased. Multiple imputation or inverse probability weighting can help, but only if you have enough observed data to justify the model. I have seen reviewers accept imputed analyses that were essentially fabrications because the missing data exceeded forty percent and the imputation model had no anchor variables. Outcome measurement needs attention to blinding of assessors, especially for subjective endpoints. Pain scores, quality of life, and functional status are all vulnerable to assessor expectations. If the outcome assessor was not blind and the study relied on patient-reported measures without calibration, the measurement domain goes to some concerns at minimum. Selective reporting requires comparing the pre-specified outcomes in the protocol with what appears in the results. I keep a simple spreadsheet with columns for domain, specific concern, evidence from text, and rating. It takes about fifteen to twenty minutes per trial once you get used to the format, compared to forty or fifty minutes if you are still learning where to look.
Get the Full Details

Common mistakes I see repeatedly
Confusing incomplete reporting with bias. Just because a trial does not mention randomization does not automatically make it high risk. The correct rating in many of those cases is some concerns, because you genuinely cannot determine the risk from available information. Some reviewers inflate their ratings to be conservative, which systematically biases the overall assessment toward finding more flawed trials than actually exist. Ignoring the granularity of the tool. RoB 2 does not give you a single summary score. The domain-level approach is the point. A trial can be low risk in randomization, some concerns in missing data, and high risk in selective reporting. Combining those into one number erases useful information and makes your review harder to interpret. Rating based on the paper's claims rather than evidence. Authors will say "allocation was concealed" in their methods section without describing how. You rate what you can verify, not what they claim. This distinction matters and most beginners do not catch it early.
When this tool does not work well
The RoB 2 tool is designed for parallel-group randomized trials with individual-level randomization. It is not suitable for cluster-randomized trials, which need the RoB 2 for clusters adaptation. It also struggles with pragmatic trials where blinding of participants and providers is intentionally absent. In those cases, labeling everything as high risk because participants knew their treatment is not useful. The tool itself acknowledges this through its signaling questions, but reviewers still routinely misapply it to non-standard designs. For non-randomized studies, you should use ROBINS-I instead. Using RoB 2 on observational data gives you a rating that looks authoritative but is structurally inappropriate. I have seen this error in published meta-analyses more often than I would like to admit. The tool also assumes you have access to the protocol or registry entry. If neither exists, your ability to assess selective reporting is severely limited, and the overall certainty of your review's findings drops accordingly. There is no workaround other than noting the limitation explicitly and discussing how it may affect interpretation.
Getting the tool
You can download the official RoB 2 tool and its accompanying handbook directly from the Cochrane website. The interactive Excel-based tool is free and updated regularly. There are also third-party implementations in RevMan and in R packages, but the official version is the most transparent about how each decision maps to the signaling questions. The handbook is not optional reading. It contains the exact decision tree for borderline cases, which is where most of your time will be spent. The tool itself is compact; the reasoning behind each rating lives in the documentation.

A note on certainty and bias
Assessing risk of bias is not the same as grading certainty of evidence. The former evaluates individual studies. The latter, usually handled through GRADE, synthesizes those evaluations into an overall confidence statement for each outcome. Do not skip the synthesis step because the bias assessment feels complete on its own. A review with all low-risk studies can still receive a low certainty rating if the effect sizes are small and the precision is wide. The two processes are linked but distinct, and treating them as interchangeable is a persistent source of confusion in the literature.