Related Experiment Video
Updated: Jul 12, 2026

05:37
An R-Based Landscape Validation of a Competing Risk Model
Published on: September 16, 2022
Evaluating large language model performance in Risk of Bias assessments: A cross-sectional validation study
Siddharth Gandhi1, Arveen Shokravi2, Yashan Chelliahpillai3
1Faculty of Medicine, Queen's University, Kingston, Ontario, Canada.
Plos One
|July 9, 2026
Summary
ChatGPT-o3 identified more high-risk randomized clinical trials (RCTs) than human reviewers, but its performance was modest. It may assist, but not replace, human assessment in systematic reviews.
Area of Science:
- Medical Informatics
- Clinical Epidemiology
- Artificial Intelligence in Healthcare
Background:
- Risk of Bias (RoB) assessment is crucial for evaluating the reliability of randomized clinical trials (RCTs).
- The Cochrane Risk of Bias 2.0 tool is a standard for RoB assessment.
- Evaluating the performance of artificial intelligence (AI) tools like ChatGPT-o3 in RoB assessment is essential.
Purpose of the Study:
- To assess the reliability and diagnostic performance of ChatGPT-o3 in conducting RoB assessments of RCTs.
- To compare ChatGPT-o3's RoB assessments with those of human reviewers using the Cochrane RoB 2.0 tool.
Main Methods:
- A methodological validation study analyzed 50 RCTs.
- ChatGPT-o3, original systematic review authors (OSRAs), and a masked human panel independently assessed RoB.
- Agreement was measured using weighted Cohen's kappa and Gwet's AC2; diagnostic performance was assessed by sensitivity, specificity, and balanced accuracy.
Main Results:
- ChatGPT-o3 identified a higher percentage of high-risk trials (34%) compared to the human panel (22%) and OSRAs (12%).
- Agreement between ChatGPT-o3 and human reviewers was modest (median kappa: 0.33 with panel; 0.14 with OSRAs).
- ChatGPT-o3's balanced accuracy for detecting high-risk trials was 0.57 and for low-risk trials was 0.66.
Conclusions:
- ChatGPT-o3 provides more conservative RoB ratings, flagging more trials as high risk than human reviewers.
- ChatGPT-o3 is currently unsuitable as a sole RoB assessor but can be a valuable adjunct tool to improve efficiency and consistency in systematic reviews.
Related Concept Videos
Bias in Epidemiological Studies
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
Reliability and Validity
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
