Related Experiment Video
Updated: Jun 19, 2026

07:13
Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
Comparing clinical decision-making between colposcopists and large language models in cervical dysplasia management:
Jan Lennart Stalp1,2, Juliane Alexandra Schneider3, Lena Steinkasserer4
1Department of Obstetrics and Gynecology, Hannover Medical School, Carl-Neuberg-Str. 1, 30625, Hannover, Germany. stalp.jan@mh-hannover.de.
Archives of Gynecology and Obstetrics
|June 18, 2026
Summary
Board-certified colposcopists and large language models (LLMs) showed similar decision-making abilities in cervical dysplasia management. LLMs excel in precancerous lesions, while clinicians manage complex cases better.
Area of Science:
- Gynecology
- Medical Artificial Intelligence
- Oncology
Background:
- Cervical dysplasia management requires complex decision-making.
- Large language models (LLMs) are emerging as potential tools in healthcare.
Purpose of the Study:
- To compare the decision-making abilities of colposcopists and LLMs (ChatGPT-4o, ChatGPT-5) in managing cervical dysplasia.
- To evaluate the performance of LLMs against a gold standard in real-life patient cases.
Main Methods:
- Prospective multicenter study using 23 anonymized patient cases with multiple-choice questions.
- Board-certified colposcopists and two LLMs answered treatment decision questions.
- Concordance rates were calculated against a gold standard defined by guideline authors.
Main Results:
- Overall concordance rates were similar: clinicians (69.6%), ChatGPT-4o (69.6%), and ChatGPT-5 (65.2%).
- ChatGPT-5 outperformed clinicians in precancerous lesions (81.8% vs. 66.4%).
- Clinicians excelled in complex cases with unspecific histopathology (86% vs. 60%) and tended to overtreat low-grade lesions.
Conclusions:
- LLMs show potential as decision support tools for straightforward cervical dysplasia cases, like precancerous lesions.
- Clinicians remain superior in managing complex or ambiguous cases.
- LLMs' practical application could be enhanced by exploring open-ended scenarios and retrieval-augmented generation.