Related Experiment Video
Updated: Jul 1, 2026

08:22
A Novel Single Animal Motor Function Tracking System Using Simple, Readily Available Software
Published on: August 31, 2018
Complementary Error Patterns Between Human Evaluators and GPT-4o in Video-Based Cardiopulmonary Resuscitation Skills
Hye Ji Park1, Daun Choi2, Choung Ah Lee1,3
1Department of Emergency Medicine, Hallym University Dongtan Sacred Heart Hospital, Hwaseong 18450, Republic of Korea.
Journal of Clinical Medicine
|June 26, 2026
Summary
GPT-4o shows promise as a decision support tool for cardiopulmonary resuscitation (CPR) skill assessment, excelling in procedural evaluations but lacking accuracy in assessing chest compression metrics. It cannot replace human experts but can help mitigate bias.
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
- Cardiopulmonary Resuscitation
Background:
- Cardiopulmonary resuscitation (CPR) skill assessment faces challenges with human evaluator subjectivity and limitations.
- Automated video-based assessment using artificial intelligence (AI) is emerging but requires validation for clinical skills.
Purpose of the Study:
- To evaluate the validity of GPT-4o for automated, video-based cardiopulmonary resuscitation (CPR) skill assessment.
- To compare GPT-4o's performance against expert human evaluators and manikin sensor data.
Main Methods:
- 130 laypersons underwent Basic Life Support (BLS) training and skill testing.
- 110 video recordings were independently assessed by expert evaluators and GPT-4o using a 12-item checklist.
- Agreement was measured using Gwet's AC1 and ICC(2,1); diagnostic accuracy was compared using McNemar's test.
Main Results:
- GPT-4o achieved near-perfect agreement (AC1 > 0.8) with experts on procedural items like calling for help.
- Agreement was poor for chest compression depth (AC1 = 0.374) and recoil (AC1 = 0.355).
- GPT-4o demonstrated high specificity but low sensitivity for compression metrics, complementing expert performance.
Conclusions:
- GPT-4o has limitations in assessing 3D spatial information from 2D videos, precluding its use as a standalone CPR evaluator.
- GPT-4o shows potential as a decision support tool to enhance CPR education by mitigating expert bias.
Keywords:
GPT-4oartificial intelligencecardiopulmonary resuscitationmedical educationskills assessmentvision-language model
