Related Experiment Video
Updated: Jan 9, 2026

The Use of Mixed Reality in Custom-Made Revision Hip Arthroplasty: A First Case Report
Published on: August 4, 2022
Retrieval-augmented ChatGPT-4o improves accuracy but reduces readability in hip arthroscopy patient education
Onur Gültekin1, Erdem Aras Sezgin2, Oğuzhan Cakır1
1Department of Orthopaedics and Traumatology, University of Health Sciences, Fatih Sultan Mehmet Training and Research Hospital, Istanbul, Türkiye.
The "deep research" mode of ChatGPT-4o offers superior accuracy and comprehensiveness for hip arthroscopy information compared to the standard model. However, the standard model provides better readability, suggesting a hybrid approach for optimal patient education.
Area of Science:
- Orthopaedic Surgery
- Artificial Intelligence in Medicine
- Health Informatics
Background:
- Large language models (LLMs) show promise for patient education but their reliability in specialized medical fields like orthopaedics is uncertain.
- Evaluating different LLM configurations is crucial for understanding their potential and limitations in providing accurate patient information.
Purpose of the Study:
- To compare the accuracy, readability, and patient-centeredness of responses from standard ChatGPT-4o and its 'deep research' mode.
- To assess the suitability of these AI models for hip arthroscopy patient education.
Main Methods:
- Thirty standardized patient questions on hip arthroscopy were used.
- Responses from standard ChatGPT-4o and its 'deep research' mode were evaluated by two orthopaedic surgeons.
- Assessments included accuracy, clarity, comprehensiveness, and readability using Likert scales and Flesch-Kincaid metrics.
Main Results:
- 'Deep research' mode demonstrated significantly higher accuracy (p=0.012) and comprehensiveness (p<0.001).
- Standard ChatGPT-4o showed superior clarity (p=0.048) and readability scores (Flesch-Kincaid Grade Level and Reading Ease Score; p<0.001).
- Readability Likert scores were comparable between models (p=0.729).
Conclusions:
- 'Deep research' mode offers greater scientific rigor, while the standard model excels in readability for hip arthroscopy patient education.
- A hybrid approach may enhance educational effectiveness, but clinical oversight is essential to prevent misinformation.
- Findings suggest modest differences and highlight the accuracy-readability trade-off in LLMs, warranting further exploratory research.
Related Concept Videos
ER Retrieval Pathway
The ER uses many checkpoints to prevent the entry of incorrectly folded or a resident protein as cargo onto a transport vesicle. These mechanisms...
Improving Translational Accuracy
Improving Translational Accuracy
Chronic Kidney Disease III: Interprofessional Care

