Related Experiment Video
Updated: Apr 8, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Comparative Analysis of Large Language Models for Pediatric Kidney Stone Patient Education: A Multi-dimensional
Fesih Ok1, Ibrahim Halil Sukur1, Zahide Orhan Ok2
1Department of Urology, Adana City Training and Research Hospital, 01370 Adana, Turkey.
Background:
Pediatric urolithiasis is an increasingly important health concern, and affected children and their families require information that is both accurate and easily understandable. Artificial intelligence (AI)-powered chatbots have become widely used sources of health information; however, the readability, quality, and reliability of their outputs remain insufficiently evaluated. This study aimed to assess the effectiveness and reliability of AI chatbots in providing patient-oriented information on pediatric kidney stone disease and to identify factors influencing the quality and readability of their responses.
Methods:
Four AI chatbots (ChatGPT-5, Google Gemini, Claude 3 Opus, and DeepSEEK) were queried with 30 standardized questions related to pediatric kidney stones. Readability was evaluated using the Average Reading Level Consensus (ARLC), Automated Readability Index (ARI), and Simple Measure of Gobbledygook (SMOG). Response quality and reliability were asssessed using the Ensuring Quality Information for Patients (EQIP) tool and Modified DISCERN score. Statistical analyses included one-way analysis of variance ANOVA, Kruskal-Wallis tests, and appropriate post hoc comparisons.
Results:
Readability differed significantly among the chatbots. Google Gemini demonstrated the highest reading levels across all metrics (ARLC: 14.93, ARI: 16.2, and SMOG: 13.32), whereas ChatGPT, Claude, and DeepSEEK produced less complex test (p < 0.001; large effect sizes, η2 = 0.195-0.512). EQIP scores did not differ significantly between models (p = 0.491, ε2 = 0.021, negligible effect), indicating comparable informational quality. In contrast, reliability varied significantly: ChatGPT and Google Gemini achieved higher Modified DISCERN scores (median 4.00) than Claude and DeepSEEK (median 3.00; p = 0.001, ε2 = 0.318, large effect). Subgroup analyses by question category revealed notable differences in performance, highlighting model-specific strenghts and limitations.
Conclusions:
Substantial variability exists in the readability and reliability of AI-generated health information on pediatric urolithiasis. Although ChatGPT and Google Gemini provided more reliable information, Google Gemini's responses were consistently more complex and less accessible. These findings emphasize the need for careful validation and language simplification of AI-generated content before its use in patient and caregiver education.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
06:28E-Patient Counseling Trial E-PACO: Computer Based Education versus Nurse Counseling for Patients to Prepare for Colonoscopy
Published on: August 1, 2019
Related Concept Videos
Urinary Tract Calculi IV: Nutrition Therapy and Prevention
Urinary Tract Calculi III: Medical Management
Urinary Tract Calculi V: Nursing Management
Chronic Kidney Disease III: Interprofessional Care
Urinary Tract Calculi VI: Surgical Management
Imaging Studies I: Kidney, Ureter, and Bladder Studies