Related Experiment Videos
Discrepancy between medical knowledge and clinical reasoning: evaluating large language models in prenatal diagnosis
Qiuling Guo1, Yanyun Du1, Jingcheng Dai2
1Department of Ultrasound, Quanzhou Maternal and Child Health Hospital (Quanzhou Children's Hospital), Quanzhou, China.
Purpose:
Large language models (LLMs) have demonstrated impressive performance across various medical examinations; however, their competency in subspecialty clinical reasoning, particularly in prenatal diagnosis, remains poorly characterized.
Materials And Methods:
A cross-sectional comparative study was conducted by using 125 questions from China's nationally standardized Prenatal Ultrasound Screening Qualification Examination. Three LLMs and two human physicians were evaluated. Questions were classified along three independent taxonomic axes: modality, disease system, and question format. Cochran's Q test was used for omnibus comparisons, with post-hoc pairwise McNemar tests where were significant.
Results:
Overall accuracy did not differ significantly among the five test-takers (p = 0.126): the chief physician scored highest (89.60%, 112/125), followed by Claude-Sonnet-4.6 (85.60%, 107/125), Grok-4 (84.80%, 106/125), ChatGPT-5 (84.00%, 105/125), and the attending physician (80.00%, 100/125). No significant differences were observed when stratified by modality (text-only, p = 0.146; imaging diagnostic questions, p = 0.603) or by any individual disease system (all p > 0.05). However, performance divergence emerged sharply across question types. In clinical vignette-based questions (n = 20), the chief physician (95.00%) significantly outperformed all other test-takers (p = 0.016), with a 25% gap over the best-performing LLM (Grok-4, 70.00%). In sequential item sets, Grok-4 (95.74%) and the chief physician (93.62%) both significantly outperformed the attending physician (76.60%; both p < 0.05). Conversely, LLMs excelled in knowledge-retrieval tasks, with ChatGPT-5 and Claude-Sonnet-4.6 both achieving 100% accuracy in matching questions. No significant differences were detected among test-takers for factual recall, matching, or multiple-choice multiple-answer questions (all p > 0.05).
Conclusion:
LLMs are well-positioned to serve as powerful knowledge augmentation tools, but are not yet ready to function as autonomous clinical reasoners in this complex, ethically consequential subspecialty.