Related Experiment Video
Updated: Jun 28, 2025

Signal Acquisition, Score Interpretation, and Economics of a Non-Invasive Point-of-Care Test for Coronary Artery Disease
Published on: August 9, 2024
ChatGPT provides inconsistent risk-stratification of patients with atraumatic chest pain
Thomas F Heston1,2, Lawrence M Lewis3
1Department of Family Medicine, University of Washington School of Medicine, Seattle, Washington, United States of America.
Background:
ChatGPT-4 is a large language model with promising healthcare applications. However, its ability to analyze complex clinical data and provide consistent results is poorly known. Compared to validated tools, this study evaluated ChatGPT-4's risk stratification of simulated patients with acute nontraumatic chest pain.
Methods:
Three datasets of simulated case studies were created: one based on the TIMI score variables, another on HEART score variables, and a third comprising 44 randomized variables related to non-traumatic chest pain presentations. ChatGPT-4 independently scored each dataset five times. Its risk scores were compared to calculated TIMI and HEART scores. A model trained on 44 clinical variables was evaluated for consistency.
Results:
ChatGPT-4 showed a high correlation with TIMI and HEART scores (r = 0.898 and 0.928, respectively), but the distribution of individual risk assessments was broad. ChatGPT-4 gave a different risk 45-48% of the time for a fixed TIMI or HEART score. On the 44-variable model, a majority of the five ChatGPT-4 models agreed on a diagnosis category only 56% of the time, and risk scores were poorly correlated (r = 0.605).
Conclusion:
While ChatGPT-4 correlates closely with established risk stratification tools regarding mean scores, its inconsistency when presented with identical patient data on separate occasions raises concerns about its reliability. The findings suggest that while large language models like ChatGPT-4 hold promise for healthcare applications, further refinement and customization are necessary, particularly in the clinical risk assessment of atraumatic chest pain patients.
More Related Videos
Related Concept Videos
Flail Chest-II
Assessment:
1. Clinical Evaluation:
History:
Flail Chest-I
Flail chest is a severe and potentially life-threatening condition characterized by the fracture of three or more adjacent ribs in multiple places. It is most commonly caused by direct impacts and trauma, such as motor vehicle accidents or injuries from a steering wheel impact. It can also occur due to falls in elderly individuals with osteoporosis, or assaults involving sharp objects.
Pathophysiology
The pathophysiology of flail chest is complex, involving fractures of...
Pneumothorax-II
Clinical Manifestations:

