Related Experiment Video
Updated: May 7, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
AI chatbots in strabismus care: A multidomain expert evaluation of caregiver-facing information
Dhiman Shweta1, Dutta Paromita2, Thacker Prolima3
1Department of Pediatric Ophthalmology & Strabismus, JPM Rotary Eye Hospital and Research Institute, Cuttack, Odisha, India-753014.
Abstract:
AimTo evaluate and compare the performance of five artificial intelligence (AI) chatbots-ChatGPT (OpenAI 4), Google Gemini, Grok (xAI), DeepSeek, and Meta Llama -in delivering accurate, clear, educational, and safe responses to caregiver-facing queries related to strabismus.MethodsSixteen standardized caregiver questions on strabismus were presented to each chatbot in independent sessions. Five fellowship-trained pediatric ophthalmologists rated each response across four domains-Accuracy, Clarity, Educational Value, and Safety-using a 5-point Likert scale (1 = Poor, 5 = Excellent). Between-chatbot differences were analyzed using cumulative link mixed models (CLMMs) with odds ratios (OR) and 95% confidence intervals (CI). Holm-adjusted pairwise contrasts corrected for multiple comparisons. Inter-rater reliability was assessed using quadratic-weighted Fleiss' κ and Gwet's AC1 to address prevalence and bias effects.ResultsChatGPT achieved the highest proportion of top ratings (≥4) for Accuracy (65%) and Clarity (59%), followed by Llama (41% and 47.5%, respectively). For Educational Value, Llama (43.8%) and Gemini (42.5%) performed slightly better, while Safety ratings were highest for Gemini (40%) and Llama (37.5%). CLMM analysis showed significant between-chatbot differences for Accuracy, Clarity, and Educational Value (p < 0.05) but not for Safety. Compared with ChatGPT, lower odds of higher ratings were seen for Grok (OR 0.48) and DeepSeek (OR 0.61). Inter-rater reliability indicated moderate agreement (Fleiss' κ = 0.59) and strong consensus (Gwet's AC1 = 0.87).ConclusionChatGPT showed superior accuracy and clarity, while Gemini and Llama excelled in educational value and safety. High expert agreement supports AI chatbots as adjuncts in pediatric ophthalmology education requiring continued validation.

