Related Experiment Video
Updated: Aug 27, 2026

Use of a Piglet Model for the Study of Anesthetic-induced Developmental Neurotoxicity (AIDN): A Translational Neuroscience Approach
Published on: June 11, 2017
GPT-4o-Powered Preanesthetic AI: Development and Validation Study
Bo-Han Lin1,2, Cheng-Wei Lu1,3, I-Ching Hou2
1Department of Anesthesiology, Far Eastern Memorial Hospital, No. 21, Section 2, Nanya S. Road, Banqiao District, New Taipei City, 220, Taiwan, +886-2-8966-7000 ext 4371.
Background:
Accurate preanesthetic assessment is essential for perioperative risk stratification, but conventional tools such as the American Society of Anesthesiologists (ASA) physical status classification and postoperative nausea and vomiting (PONV) risk scores may be affected by subjective judgment, incomplete documentation, and fragmented clinical data. Large language models may support preanesthetic assessment by integrating structured and unstructured clinical information.
Objective:
This study aimed to develop and retrospectively validate a GPT-4o (OpenAI)-powered AI system for preanesthetic assessment. The primary validation focus was the agreement between AI-generated and clinician-assigned ASA physical status classifications. PONV risk stratification was evaluated as an additional clinically relevant performance outcome. A secondary objective was to assess the incremental contribution of National Health Insurance (NHI) cloud data to model performance.
Methods:
This single-center retrospective validation study included 600 adult surgical patients randomly sampled from 4404 eligible inpatient surgical patients between January and May 2025. The system processed structured and unstructured electronic health record data, with and without NHI cloud data, to generate preanesthetic assessment outputs. ASA agreement was assessed using Cohen κ with 95% CIs. PONV prediction was evaluated using sensitivity, specificity, overall accuracy, positive predictive value, negative predictive value, confusion matrix analysis, the 2-sided Fisher-Freeman-Halton exact test, and Cramér V.
Results:
Among the 600 patients, clinician-assigned ASA classifications were ASA I in 39 patients, ASA II in 498 patients, ASA III in 57 patients, and ASA IV in 6 patients. Agreement between AI-generated and clinician-assigned ASA classifications was high when NHI data were incorporated (κ=0.883, 95% CI 0.841-0.946), whereas agreement was lower without NHI data (κ=0.518, 95% CI 0.412-0.691). A total of 49 (8.2%) patients experienced documented PONV within 24 hours after surgery. Using the high-risk category as the primary test-positive threshold, the model achieved a sensitivity of 34.7% (95% CI 22.9-48.7), a specificity of 99.1% (95% CI 97.9-99.6), a positive predictive value of 77.3%, and a negative predictive value of 94.5%. The association between the model-assigned PONV risk categories and observed PONV outcomes was statistically significant (P<.001, Cramér V=0.531).
Conclusions:
The GPT-4o-powered preanesthetic assessment system demonstrated feasibility and high agreement with clinician-assigned ASA classifications when NHI cloud data were incorporated. For PONV risk stratification, the system showed high specificity and a high negative predictive value but modest sensitivity, indicating that it identified a small high-risk subgroup with a high observed PONV incidence but did not capture all patients who developed PONV. These findings support the use of large language model-based decision-support tools as adjuncts to, rather than replacements for, anesthesiologist judgment.