Related Experiment Video
Updated: Jan 26, 2026

Isokinetic Robotic Device to Improve Test-Retest and Inter-Rater Reliability for Stretch Reflex Measurements in Stroke Patients with Spasticity
Published on: June 12, 2019
Risk of bias in nonrandomized studies of interventions showed low inter-rater reliability and challenges in its
Silvia Minozzi1, Michela Cinquini2, Silvia Gianola3
1Department of Epidemiology, Lazio Regional Health Service, via Cristoforo Colombo 112, 00147 Rome, Italy; Department of Biomedical Sciences for Health, University of Milan, via Carlo Pascal 36, 20133 Milan, Italy.
Objective:
To assess the inter-rater reliability (IRR) and usability of the risk of bias in nonrandomized studies of interventions tool (ROBINS-I).
Study Design And Setting:
We designed a cross-sectional study. Five raters independently applied ROBINS-I to the nonrandomized cohort studies in three systematic reviews on vaccines, opiate abuse, and rehabilitation. We calculated Fleiss' Kappa for multiple raters as a measure of IRR and discussed the application of ROBINS-I to identify difficulties and possible reasons for disagreement.
Results:
Thirty one studies were included (195 evaluations). IRRs were slight for overall judgment (IRR 0.06, 95% CI 0.001 to 0.12) and individual domains (from 0.04, 95% CI -0.04 to 0.12 for the domain "selection of reported results" to 0.18, 95% CI 0.10 to 0.26 for the domain "deviation from intended interventions"). Mean time to apply the tool was 27.8 minutes (SD 12.6) per study. The main difficulties were due to poor reporting of primary studies, misunderstanding of the question, translation of questions into a final judgment, and incomplete guidance.
Conclusion:
We found ROBINS-I difficult and demanding, even for raters with substantial expertise in systematic reviews. Calibration exercises and intensive training before its application are needed to improve reliability.
Insights
The Risk of Bias In Nonrandomized Studies of Interventions (ROBINS-I) tool showed slight inter-rater reliability (IRR) and was found difficult to use. Intensive training is needed to improve IRR for this risk of bias assessment tool.
Area of Science:
- Evidence-based medicine
- Systematic reviews
- Health research methodology
Background:
- Assessing the risk of bias in nonrandomized studies of interventions is crucial for evidence synthesis.
- The Risk of Bias In Nonrandomized Studies of Interventions (ROBINS-I) tool was developed to standardize this assessment.
- Evaluating the usability and reliability of the ROBINS-I tool is essential for its effective implementation.
Purpose of the Study:
- To evaluate the inter-rater reliability (IRR) of the ROBINS-I tool.
- To assess the usability of the ROBINS-I tool among experienced raters.
- To identify challenges and reasons for disagreement in applying the ROBINS-I tool.
Main Methods:
- A cross-sectional study design was employed.
- Five independent raters applied the ROBINS-I tool to 31 nonrandomized cohort studies from three systematic reviews.
- Inter-rater reliability was measured using Fleiss' Kappa, and time taken to apply the tool was recorded.
Main Results:
- Overall IRR for the ROBINS-I tool was slight (IRR 0.06).
- IRR varied across domains, with the lowest for 'selection of reported results' (IRR 0.04) and highest for 'deviation from intended interventions' (IRR 0.18).
- Raters reported difficulties due to poor primary study reporting, misunderstanding questions, and incomplete guidance, with an average application time of 27.8 minutes per study.
Conclusions:
- The ROBINS-I tool was found to be difficult and demanding, even for experienced systematic reviewers.
- Calibration exercises and intensive training are recommended prior to using the ROBINS-I tool to enhance reliability.
- Improvements in reporting of primary studies may also enhance the usability and reliability of the ROBINS-I tool.
Related Concept Videos
Bias in Epidemiological Studies
Confirmation Biases
Hindsight Biases
Reliability and Validity
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Nursing Interventions I: Taxonomy of Nursing Interventions
A nursing intervention is a treatment or action based on scientific concepts and knowledge from the nursing, behavioral, and physical sciences. Identifying and prioritizing nursing interventions based on the desired outcome...

