Related Experiment Video
Updated: Sep 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Are artificial intelligence systems ready for pediatric surgical decision-making? A comparative evaluation of large
1Department of Pediatric Surgery, Faculty of Medicine, Hitit University, Çorum, Turkey. mehmet-mtn@hotmail.com.
Purpose:
This study aimed to compare the diagnostic, investigative, and treatment decisions generated by five contemporary large language models (LLMs) across 50 pediatric surgery clinical scenarios against a reference standard established by pediatric surgery experts.
Methods:
Fifty guideline-based clinical scenarios representing a broad range of pediatric surgical subspecialties were developed. Diagnosis, investigation, and treatment decisions were assessed for each scenario. A consensus reference standard was established by two pediatric surgeons. ChatGPT (GPT-5.5), Claude (Sonnet 4.6), Gemini (3.5 Flash), DeepSeek-V3, and Perplexity (1.0) were evaluated independently using standardized prompts. Accuracy, observed agreement (Po), Gwet's AC1 with 95% confidence intervals, and McNemar tests were used to evaluate model performance and agreement with the pediatric surgeon.
Results:
Overall, the pediatric surgeon correctly classified 144 of 150 decisions (Po = 0.960). Claude achieved the highest overall performance (145/150; Po = 0.967), followed by ChatGPT and Gemini (144/150 each; Po = 0.960). DeepSeek-V3 and Perplexity achieved Po values of 0.913 and 0.873, respectively. Claude achieved perfect diagnostic performance (50/50; Po = 1.000). ChatGPT and Gemini performed best for investigations (48/50; Po = 0.960), while Claude, ChatGPT, and Gemini each achieved 48/50 correct treatment decisions (Po = 0.960). No significant differences were identified between the pediatric surgeon and any AI model with respect to diagnostic, investigative, or treatment performance (all p > 0.05). In the overall analysis, only Perplexity demonstrated a performance level significantly different from that of the pediatric surgeon (p = 0.004).
Conclusions:
LLMs showed high accuracy and strong agreement with expert judgment under standardized guideline-based clinical scenarios, particularly Claude, ChatGPT, and Gemini. The clustering of errors within the investigation and treatment domains suggests that, although these models show considerable promise as clinical decision-support tools, they should be used under clinician supervision rather than as independent decision-makers.
Clinical Trial Registration:
Not applicable.