Related Experiment Video
Updated: Apr 10, 2026

Estimate the Cognitive Load Using Electrocardiographic Measure: A Human-AI Collaborative Task
Published on: December 5, 2025
A new peer reviewer? Comparing AI with human performance in randomized controlled trial risk-of-bias assessment
Jonathan Lettner1,2, Marko Ostojic3, Aleksandra Królikowska4,5
1Faculty of Health Sciences Brandenburg, Brandenburg Medical School Theodor Fontane, Brandenburg an der Havel, Germany.
Background:
Risk-of-bias (RoB) assessment is essential for evidence synthesis but remains time-consuming and inherently subjective. Artificial intelligence (AI) may improve the efficiency of systematic reviews; however, its reliability in reproducing expert RoB judgements remains uncertain.
Objectives:
To compare the performance of AI models and human raters in RoB assessment of randomized controlled trials (RCTs) using the revised JBI critical appraisal tool.
Material And Methods:
Thirteen RCTs published between 2023 and 2025 in orthopedic journals were independently assessed by 2 human raters (an expert (R1) and a novice (R2)) and 2 AI models (ChatGPT-4.0 (CGPT) and DeepSeek-R1 (DS)) using the 13-domain JBI checklist. Deep-reasoning functionalities (e.g., chain-of-thought prompting) were applied. Inter-rater agreement, deviations from the expert assessment (reference standard), and binary disagreements (e.g., Yes vs No) were analyzed to evaluate consistency.
Results:
The AI models demonstrated high inter-model agreement (91%), exceeding human-AI agreement (CGPT vs R1: 64%; DS vs R1: 68%). However, both AI systems showed substantial divergence from expert judgements in interpretive domains, including allocation concealment (Q2), blinding (Q7), and overall trial design (Q13), with deviation rates ranging from 30% to 38.5%. Binary decision reversals were more frequent in AI assessments (CGPT: 8.9%; DS: 7.7%) than in the human comparison (R2 vs R1: 2.4%). Human raters showed stronger agreement in contextual interpretation (R1-R2: 89.3%), whereas AI models performed better in rule-based domains (Q8/Q9: 100% agreement).
Conclusions:
AI can reliably support the automation of objective components of RoB assessment but remains limited in handling interpretive, context-dependent judgements. A hybrid approach combining AI-assisted pre-screening with expert evaluation may enhance the scalability of systematic reviews without compromising methodological rigor.
More Related Videos
07:34Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
06:02Evaluating Usability Aspects of a Mixed Reality Solution for Immersive Analytics in Industry 4.0 Scenarios
Published on: October 6, 2020
Related Concept Videos
Blinding
Randomized Experiments
Simple randomization
Simple...
Blind Procedures
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Bias in Epidemiological Studies
Regression Toward the Mean