Related Experiment Video
Updated: Jun 20, 2026

Movement Retraining using Real-time Feedback of Performance
Published on: January 17, 2013
ChatGPT is a comprehensive education tool for patients with patellar tendinopathy, but it currently lacks accuracy
Jie Deng1, Lun Li2, Jelle J Oosterhof3
1Department of Orthopedics and Sports Medicine, Erasmus MC University Medical Center, Rotterdam, the Netherlands; Department of Radiology and Nuclear Medicine, Erasmus MC University Medical Center, Rotterdam, the Netherlands.
Background:
Generative artificial intelligence tools, such as ChatGPT, are becoming increasingly integrated into daily life, and patients might turn to this tool to seek medical information.
Objective:
To evaluate the performance of ChatGPT-4 in responding to patient-centered queries for patellar tendinopathy (PT).
Methods:
Forty-eight patient-centered queries were collected from online sources, PT patients, and experts and were then submitted to ChatGPT-4. Three board-certified experts independently assessed the accuracy and comprehensiveness of the responses. Readability was measured using the Flesch-Kincaid Grade Level (FKGL: higher scores indicate a higher grade reading level). The Patient Education Materials Assessment Tool (PEMAT) evaluated understandability, and actionability (0-100%, higher scores indicate information with clearer messages and more identifiable actions). Semantic Textual Similarity (STS score, 0-1; higher scores indicate higher similarity) assessed variation in the meaning of texts over two months (including ChatGPT-4o) and for different terminologies related to PT.
Results:
Sixteen (33%) of the 48 responses were rated accurate, while 36 (75%) were rated comprehensive. Only 17% of treatment-related questions received accurate responses. Most responses were written at a college reading level (median and interquartile range [IQR] of FKGL score: 15.4 [14.4-16.6]). The median of PEMAT for understandability was 83% (IQR: 70%-92%), and for actionability, it was 60% (IQR: 40%-60%). The medians of STS scores in the meaning of texts over two months and across terminologies were all ≥ 0.9.
Conclusions:
ChatGPT-4 provided generally comprehensive information in response to patient-centered queries but lacked accuracy and was difficult to read for individuals below a college reading level.
Related Concept Videos
Therapeutic Communication
Verbal communication depends on language or a prescribed way of using words so that people can share information effectively. The critical aspects of verbal...
Techniques of therapeutic communication I: Active Listening, Sharing Observations, Validation, and Using Touch
Therapeutic communication is not the same as social interaction. Social interaction has no goal or purpose and consists of casual information sharing, whereas therapeutic communication has a plan or purpose for the conversation. Therapeutic...
Methods of Documentation III: PIE
Myasthenia Gravis: Diagnostic Tests
The edrophonium test is a diagnostic tool for myasthenia gravis. It involves...
Pulmonary Function Tests
Pulmonary Function Tests are crucial diagnostic tools for assessing respiratory function, particularly in patients with chronic respiratory disorders. They comprehensively evaluate lung volumes, ventilatory function, breathing mechanics, diffusion, and gas exchange. These tests help diagnose pulmonary diseases and play a significant role in monitoring disease progression, evaluating disability, and assessing response to therapy.
PFTs involve using a spirometer, a...
Assessment of the Gastrointestinal System II: Health Perception Pattern
Health Perception Patterns
Health perception patterns offer valuable insights into a patient's lifestyle habits and how they may impact their GI health. These patterns include:

