Related Experiment Video
Updated: May 24, 2025

A Rat Tibial Growth Plate Injury Model to Characterize Repair Mechanisms and Evaluate Growth Plate Regeneration Strategies
Published on: July 4, 2017
Evaluating the Quality and Readability of Generative Artificial Intelligence (AI) Chatbot Responses in the Management
Christopher E Collins1, Peter A Giammanco2,1, Monica Guirgus1
1Orthopedic Surgery, California University of Science and Medicine, Colton, USA.
Introduction:
The rise of artificial intelligence (AI), including generative chatbots like ChatGPT (OpenAI, San Francisco, CA, USA), has revolutionized many fields, including healthcare. Patients have gained the ability to prompt chatbots to generate purportedly accurate and individualized healthcare content. This study analyzed the readability and quality of answers to Achilles tendon rupture questions from six generative AI chatbots to evaluate and distinguish their potential as patient education resources.
Methods:
The six AI models used were ChatGPT 3.5, ChatGPT 4, Gemini 1.0 (previously Bard; Google, Mountain View, CA, USA), Gemini 1.5 Pro, Claude (Anthropic, San Francisco, CA, USA) and Grok (xAI, Palo Alto, CA, USA) without prior prompting. Each was asked 10 common patient questions about Achilles tendon rupture, determined by five orthopaedic surgeons. The readability of generative responses was measured using Flesch-Kincaid Reading Grade Level, Gunning Fog, and SMOG (Simple Measure of Gobbledygook). The response quality was subsequently graded using the DISCERN criteria by five blinded orthopaedic surgeons.
Results:
Gemini 1.0 generated statistically significant differences in ease of readability (closest to average American reading level) than responses from ChatGPT 3.5, ChatGPT 4, and Claude. Additionally, mean DISCERN scores demonstrated significantly higher quality of responses from Gemini 1.0 (63.0±5.1) and ChatGPT 4 (63.8±6.2) than ChatGPT 3.5 (53.8±3.8), Claude (55.0±3.8), and Grok (54.2±4.8). However, the overall quality (question 16, DISCERN) of each model was averaged and graded at an above-average level (range, 3.4-4.4).
Discussion And Conclusion:
Our results indicate that generative chatbots can potentially serve as patient education resources alongside physicians. Although some models lacked sufficient content, each performed above average in overall quality. With the lowest readability and highest DISCERN scores, Gemini 1.0 outperformed ChatGPT, Claude, and Grok and potentially emerged as the simplest and most reliable generative chatbot regarding management of Achilles tendon rupture.
More Related Videos
04:49Author Spotlight: Enhancing Post-Stroke Upper Limb Rehabilitation with Robotic Technologies for Improved Motor Recovery and Functional Outcomes
Published on: September 6, 2024
08:19Author Spotlight: Unraveling the Mechanobiology of Tendon Impingement – A Multiaxial Murine Hind Limb Explant Model
Published on: December 8, 2023
Related Concept Videos
Directly Acting Muscle Relaxants: Dantrolene and Botulinum Toxin
The binding of dantrolene to the RYR1...
Fractures: Bone Repair
Minor fractures with no bone displacement are treated by immobilizing the fractured bone using a cast or splint. However, in the case of fractures with displaced bones, the broken bones are repositioned before immobilization to ensure successful healing without deformation and loss of function. The realignment of fractured bone ends is performed through a process called reduction. If the...
Ankle Joint