Related Experiment Video
Updated: Apr 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Evaluating Artificial Intelligence Conversational Platforms for Parental Queries on Antenatal Hydronephrosis: A
Nidhi Patil1, Anjan Kumar Dhua1, Bitesh Kumar1
1Department of Pediatric Surgery, All India Institute of Medical Sciences, New Delhi, India.
Background:
Antenatal hydronephrosis is the most common fetal anomaly detected on routine ultrasound. Parents often seek immediate information online and turn to artificial intelligence (AI) conversational platforms, which are easily available on mobile. The accuracy and reliability of these responses in sensitive pediatric surgical contexts are unknown.
Objectives:
To compare the quality of responses generated by ChatGPT, Gemini, and Claude to standardized parent-relevant questions on antenatal hydronephrosis.
Materials And Methods:
Five key questions were developed and vetted by three independent pediatric surgeons. To minimize bias, an independent person with a nonmedical background posed these questions to ChatGPT, Gemini, and Claude using a standardized background scenario. Responses were documented verbatim. Three pediatric surgeons then independently assessed each response for veracity, clarity, comprehensiveness, tone, and safety, using a 5-point Likert scale. Scores were analyzed using descriptive statistics, analysis of variance (ANOVA)/Kruskal-Wallis for inter-platform comparison, and intraclass correlation (ICC) for inter-rater reliability.
Results:
Across five assessment parameters, mean scores for the three platforms ranged between 3.53 and 4.13 on a 5-point scale. No statistically significant inter-platform differences were identified by ANOVA or Kruskal-Wallis tests (all P > 0.05). Inter-rater reliability was limited, with ICC (2, k) values ranging from 0.00 (poor) to 0.60 (moderate), indicating variability in expert interpretation of AI-generated responses.
Conclusions:
The three AI conversational platforms produced broadly comparable outputs. While responses were generally clear, reassuring, and safe, the poor-to-moderate inter-rater agreement underscores heterogeneity in expert appraisal. These findings highlight that AI platforms should be considered adjuncts rather than substitutes for professional counseling, and their evolving nature warrants ongoing evaluation.

