Related Experiment Video
Updated: Sep 27, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Flatfoot and artificial intelligence: Can language models be trusted for patient education?
Hasan Emirhan Usta1, Ruhat Ünlü2
1Department of Orthopedics and Traumatology, Gebze Fatih State Hospital, Kocaeli, Türkiye.
Abstract:
Flatfoot (pes planus) is common in children and adults. Although often asymptomatic, it may alter lower-limb biomechanics and contribute to pain or injury. Patients and caregivers increasingly use artificial intelligence chatbots such as chat generative pretrained transformer (ChatGPT) and Google Gemini for medical information, yet their responses on flatfoot remain insufficiently studied. This study compared the factual accuracy, added value, omissions, and readability of responses generated by ChatGPT and Google Gemini to standardized patient-oriented questions. In this cross-sectional comparative study, 15 standardized patient-oriented questions were developed from clinical guidelines, peer-reviewed literature, and educational resources from established orthopedic societies, then refined by 2 orthopedic surgeons. Each question was submitted to ChatGPT (OpenAI) and Google Gemini (Google LLC) in July 2025 using identical wording in independent chat sessions. Responses were anonymized and independently rated by 2 board-certified orthopedic surgeons for factual accuracy, added value, and omissions using an investigator-developed 5-point rubric. Discrepant assessments were reviewed by a 3rd senior orthopedic surgeon. Primary analyses used mean scores of the 2 reviewers. Readability was assessed using the Flesch-Kincaid grade level and Flesch reading ease score. Paired differences were analyzed using the Wilcoxon signed-rank test, with P < .05 considered significant. Both models produced generally accurate responses, with mean factual accuracy scores above 4/5. Gemini scored higher than ChatGPT for factual accuracy (4.7 ± 0.2 vs 4.2 ± 0.3, P = .01), added value (4.6 ± 0.3 vs 4.0 ± 0.4, P = .02), and omissions (4.8 ± 0.2 vs 3.9 ± 0.4, P < .001), with higher scores indicating fewer clinically relevant omissions. Gemini also generated longer responses (165 ± 20 vs 120 ± 15 words, P < .001) and showed better readability, with lower Flesch-Kincaid grade level (9.2 ± 1.1 vs 11.8 ± 1.4, P < .001) and higher Flesch reading ease score (59.4 ± 6.2 vs 42.5 ± 5.3, P < .001). Both models generated generally accurate answers to standardized flatfoot questions. Gemini performed better across evaluation domains and readability measures. As these findings are time-specific, repeated assessments and direct patient and caregiver evaluations are needed before broader patient-facing use can be recommended.