Related Experiment Video
Updated: May 26, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
Can ChatGPT Take My Call? evaluating AI in gynecologic oncology telephone triage
Ann Marie Mercier1, Sharonne Holtzman1, Aashna Saini1
1Department of Obstetrics, Gynecology, and Reproductive Science, Division of Gynecologic Oncology, Icahn School of Medicine at Mount Sinai, United States.
Background:
As access to artificial intelligence (AI) expands, patients and clinicians increasingly rely on these platforms for medical information and guidance. This study assessed ChatGPT's ability to triage and respond to gynecologic oncology (GO) patient telephone calls.
Methods:
In this cross-sectional study, 30 patient scenarios received by on-call GO fellows were evaluated using ChatGPT-4o. Four physicians independently rated responses for accuracy, comprehensiveness, clarity, relevance, and applicability using a 5-point Likert scale (1 = very poor; 5 = very good) and assessed the presence of misinformation (yes/no). Wilcoxon signed-rank tests compared mean ratings with a predefined threshold of 3.0 ("acceptable"), with significance set at p < 0.05. The prevalence of misinformation was calculated with 95% confidence intervals (CI). Agreement with provider recommendations was assessed using Cohen's kappa. Misclassification patterns (over-triage vs under-triage) were summarized.
Results:
Among 30 scenarios, exact agreement between AI and fellows occurred in 26 cases (86.7%). Misinformation was identified in 4 responses (3.3%; 95% CI, 0.6-16.7%). Misclassifications included three cases (10.0%) over-triaged to clinic within one week and one case (3.3%) under-triaged to stay home rather than be seen within one week. In all 10 scenarios where the fellows' recommended emergency department evaluation, AI issued the same recommendation. Agreement between AI and fellow triage decisions was high (κ = 0.87). Across all categories, physician ratings of AI responses significantly exceeded the acceptable threshold (p < 0.001).
Conclusion:
ChatGPT demonstrated high reliability in triaging and responding to GO patient telephone calls, producing clear, accurate, and clinically appropriate guidance. Its consistent performance in urgent scenarios and tendency toward conservative triage suggest that AI may serve as a valuable adjunct to support on-call personnel and enhance after-hours triage workflows in GO.
