Related Experiment Video
Updated: Sep 11, 2025

05:01
A Detailed Protocol for Physiological Parameters Acquisition and Analysis in Neurosurgical Critical Patients
Published on: October 17, 2017
7.1K
Post-deployment Monitoring of AI Performance in Intracranial Hemorrhage Detection by ChatGPT
Eric Rohren1, Mohadese Ahmadzade2, Sofia Colella3
1Department of Radiology, Baylor College of Medicine, Houston, Texas (E.R.).
Academic Radiology
|August 12, 2025
Summary
This study evaluated an AI system for detecting intracranial hemorrhage (ICH) and found ChatGPT-4 Turbo useful for monitoring AI performance. Continuous AI monitoring is crucial for reliable clinical deployment.
Area of Science:
- Radiology
- Artificial Intelligence
- Medical Informatics
Background:
- Artificial intelligence (AI) systems are increasingly used in radiology for tasks like intracranial hemorrhage (ICH) detection.
- Continuous monitoring of AI performance post-deployment is essential for ensuring diagnostic accuracy and patient safety.
- Large language models (LLMs) like ChatGPT-4 Turbo offer potential for automated analysis of clinical data.
Purpose of the Study:
- To evaluate the real-world performance of the Aidoc AI system for detecting intracranial hemorrhage (ICH) in head CT examinations.
- To assess the effectiveness of ChatGPT-4 Turbo in monitoring the performance of the Aidoc AI system.
- To identify factors influencing the false positive rate of the Aidoc AI system.
Main Methods:
- A retrospective analysis of 332,809 head CT examinations from 37 US radiology practices (December 2023-May 2024).
- Utilized a HIPAA-compliant ChatGPT-4 Turbo to extract data from radiology reports for AI monitoring.
- Established ground truth through radiologist review of 200 selected cases and calculated performance metrics for AI and radiologists.
Main Results:
- ChatGPT-4 Turbo achieved high accuracy (PPV 1, NPV 0.988, AUC 0.996) in identifying ICH from reports.
- Aidoc's false positives were associated with specific scanner manufacturers (Philips) and artifacts, but reduced by midline shift and mass effect.
- Aidoc-assisted radiologists demonstrated high sensitivity (0.936) and specificity (1).
Conclusions:
- Continuous performance monitoring of clinical AI systems is vital.
- LLMs like ChatGPT-4 Turbo provide a scalable method for evaluating AI performance.
- Automated monitoring enhances the reliability of AI deployment in diagnostic workflows.

