Related Experiment Video
Updated: Aug 10, 2026

03:49
Anogenital Distance and Perineal Measurements of the Pelvic Organ Prolapse (POP) Quantification System
Published on: September 20, 2018
Performance of 5 Large Language Models in Perioperative Consultation for Pediatric Hypospadias: Cross-Sectional
Ting Kang1, Chi Yuan1, Xinyu Hu1
1Department of Pediatric Surgery, West China Hospital of Sichuan University, 37 Guoxue Xiang, Wuhou District, Chengdu, Sichuan, China, 86 189-8060-6946.
Journal of Medical Internet Research
|July 29, 2026
Summary
Large language models (LLMs) show varied performance in pediatric hypospadias care. Gemini-2.5-Pro ranked highest overall, but high citation accuracy does not ensure clinical safety for AI in medicine.
Area of Science:
- Artificial Intelligence in Medicine
- Pediatric Urology
- Health Informatics
Background:
- Hypospadias is a common congenital condition requiring surgical intervention.
- Caregivers need comprehensive perioperative information, presenting a challenge for health education.
- Large language models (LLMs) offer potential for patient education but require evaluation in specialized fields like pediatric urology.
Purpose of the Study:
- To evaluate the performance of five LLMs in providing perioperative information for pediatric hypospadias.
- To identify key dimensions prioritized by clinicians and caregivers when assessing AI-generated medical information.
- To analyze the relationship between citation accuracy and clinical safety of LLM responses.
Main Methods:
- A cross-sectional study evaluated 5 LLMs using 10 high-priority questions on pediatric hypospadias.
- Responses were assessed by 23 pediatric urology experts and 36 primary caregivers using double-blind forced-ranking.
- Citation authenticity was verified, and clinical safety was audited by specialists using a severity scale.
Main Results:
- Gemini-2.5-Pro ranked first overall, favored by both experts and caregivers.
- Citation accuracy varied significantly, with OpenEvidence showing high verifiability but DeepSeek and Zhipu Qingyan exhibiting high fabrication rates.
- A clinical safety audit identified 9 Severe-level safety flags, with OpenEvidence having the highest burden, while Gemini-2.5-Pro had the lowest.
Conclusions:
- High citation accuracy does not guarantee clinical safety in LLM-generated medical content.
- No single LLM excelled across all evaluated dimensions; Gemini-2.5-Pro was comprehensive yet relied on nonspecific guidelines.
- A tiered human-machine collaboration model with clinician oversight is recommended for AI in high-risk perioperative care.

