Related Experiment Video
Updated: Jun 27, 2025

Assessment and Evaluation of the High Risk Neonate: The NICU Network Neurobehavioral Scale
Published on: August 25, 2014
An evaluation of the capabilities of language models and nurses in providing neonatal clinical decision support
Chedva Levin1, Tehilla Kagan2, Shani Rosen3
1Faculty of School of Life and Health Sciences, Nursing Department, The Jerusalem College of Technology-Lev Academic Center, Jerusalem, Israel; The Department of Vascular Surgery, The Chaim Sheba Medical Center, Tel Hashomer, Ramat Gan, Tel Aviv, Israel.
Aim:
To assess the clinical reasoning capabilities of two large language models, ChatGPT-4 and Claude-2.0, compared to those of neonatal nurses during neonatal care scenarios.
Design:
A cross-sectional study with a comparative evaluation using a survey instrument that included six neonatal intensive care unit clinical scenarios.
Participants:
32 neonatal intensive care nurses with 5-10 years of experience working in the neonatal intensive care units of three medical centers.
Methods:
Participants responded to 6 written clinical scenarios. Simultaneously, we asked ChatGPT-4 and Claude-2.0 to provide initial assessments and treatment recommendations for the same scenarios. The responses from ChatGPT-4 and Claude-2.0 were then scored by certified neonatal nurse practitioners for accuracy, completeness, and response time.
Results:
Both models demonstrated capabilities in clinical reasoning for neonatal care, with Claude-2.0 significantly outperforming ChatGPT-4 in clinical accuracy and speed. However, limitations were identified across the cases in diagnostic precision, treatment specificity, and response lag.
Conclusions:
While showing promise, current limitations reinforce the need for deep refinement before ChatGPT-4 and Claude-2.0 can be considered for integration into clinical practice. Additional validation of these tools is important to safely leverage this Artificial Intelligence technology for enhancing clinical decision-making.
Impact:
The study provides an understanding of the reasoning accuracy of new Artificial Intelligence models in neonatal clinical care. The current accuracy gaps of ChatGPT-4 and Claude-2.0 need to be addressed prior to clinical usage.
Related Concept Videos
Patient-centered Care
Current Trends in Nursing II
Nursing Interventions II: Selecting and Classifying the Nursing Interventions
Critical Thinking I
Role of Communication in the Nursing Process I: Assessment and Diagnosis
The nursing process considers the patient's emotional and physical well-being. The process can be repeated or stopped at any point if judged essential. Assessment is the first step in the nursing...
Nursing Evaluation

