Related Experiment Video
Updated: Jan 17, 2026

A Clinical Trial Assessing the Safety, Efficacy, and Delivery of Olive-Oil-Based Three-Chamber Bags for Parenteral Nutrition
Published on: September 20, 2019
ChatGPT-4 in Nursing Research: A Methodological Evaluation of Bias Risk in Randomized Controlled Trials
Metin Tuncer1, Gülsüm Zekiye Tuncer2
1Department of Nursing Fundamentals, Gümüşhane University, Gümüşhane, Turkey.
Background:
Conducting bias assessments in systematic reviews is a time-consuming process that involves subjective judgments. The use of artificial intelligence (AI) technologies to perform these assessments can potentially save time and enhance consistency. Nevertheless, the efficacy of AI technologies in conducting bias assessments remains inadequately explored.
Aim:
This study aims to evaluate the efficacy of ChatGPT-4o in assessing bias using the revised Cochrane RoB2 tool, focusing on randomized controlled trials in nursing.
Methods:
ChatGPT-4o was provided with the RoB2 assessment guide in the form of a PDF document and instructed to perform bias assessments for the 80 open-access RCTs included in the study. The results of the bias assessments conducted by ChatGPT-4o for each domain were then compared with those of the meta-analysis authors using Cohen's weighted kappa analysis.
Results:
Weighted Cohen's kappa values showed better agreement in bias in the measurement of the outcome (D4, 0.22) and bias arising from the randomization process (D1, 0.20), while negative values in bias due to missing outcome data (D3, -0.12) and bias in the selection of the reported result (D5, -0.09) indicated poor agreement. The highest accuracy was observed in D5 (0.81), and the lowest in D1 (0.60). F1 scores were highest in bias due to deviations from intended interventions (D2, 0.74) and lowest in D3 (0.00) and D5 (0.00). Specificity was higher in D5 (0.93) and D3 (0.82), while sensitivity and precision were low in these domains.
Conclusions:
The agreement between ChatGPT-4o and the meta-analysis studies in the same RCT assessments is generally low. This indicates that ChatGPT-4o requires substantial enhancements before it can be used as a reliable tool for bias risk assessments.
Clinical Relevance:
The AI-based tools have the potential to expedite bias assessment in systematic reviews. However, this study demonstrates that ChatGPT-4o, in its current form, lacks sufficient consistency, indicating that such tools should be integrated cautiously and used under continuous human oversight, particularly in evidence-based evaluations that inform clinical decision-making.
Related Concept Videos
Bias in Epidemiological Studies
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast,...
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
The Scientific Method in Nursing Process
When using research findings to change practice, one must understand the process used to guide a study. The scientific method is a systematic, step-by-step process that supports the data's validity, reliability, and generalizability. As a result, findings can be...
Randomized Experiments
Simple randomization
Simple...
Blinding
