Related Experiment Video
Updated: Jan 15, 2026

The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve
Published on: September 7, 2022
Can Patients Differentiate Responses by Orthopaedic Surgeons and ChatGPT Regarding Total Hip Arthroplasty?
Anoop K Prasad1, Perry L Lim, Amy Z Blackburn
1From the Department of Orthopaedic Surgery, Chelsea and Westminster Hospital. London, England, UK (Prasad), the Department of Orthopaedic Surgery, Massachusetts General Hospital, Harvard Medical School, Boston, MA (Lim, Cote, Alpaugh, Melnic, and Bedair), the Department of Orthopaedic Surgery, Newton-Wellesley Hospital, Newton, MA (Lim, Cote, Alpaugh, Melnic, and Bedair), the Department of Orthopaedic Surgery, University of California Los Angeles, Los Angeles, CA (Blackburn), the Department of Orthopaedic Surgery, Royal London Hospital for Integrated Medicine, Orthopaedics, London, England, UK (Ensor), the Department of Radiology, Massachusetts General Hospital, Boston, MA (Succi), and the General Medicine Division, Massachusetts General Hospital, Boston, MA (Sepucha).
Introduction:
The digital age offers vast health information, boosting patient health literacy but also fostering challenges like misinformation. Large language models-driven chatbots such as chat generative pretrained transformer 3 (ChatGPT 3) blur human artificial intelligence (AI) content creation boundaries, potentially affecting healthcare literacy. This study aimed to assess perceived message credibility by patients when comparing healthcare descriptions by ChatGPT 3 or orthopaedic attendings.
Methods:
A cross-sectional survey of 160 patients assessed hip arthroplasty topics at increasing complexities: (1) "What is hip arthritis?"; (2) "What is a total hip replacement?"; (3) "What are the risks of undergoing a hip replacement?"; and (4) "What is the treatment of a periprosthetic joint infection of the hip?." Patients received one question randomly and compared four answers: one from ChatGPT and three from orthopaedic surgeons. Patients rated each answer (1 to 7) on credibility (accuracy, authenticity, believability) and indicated their most trusted answer. Multilevel modeling accounted for varying intercepts, providing estimated scores for each source.
Results:
Across all four questions, ChatGPT matched or exceeded the performance of orthopaedic surgeons across the domains of accuracy, authenticity, and believability. In three questions, ChatGPT surpassed at least one surgeon in one domain and, in two questions, outperformed at least two surgeons in two or more domains. Multilevel modeling revealed that ChatGPT scored the highest across all three credibility domains. Notably, ChatGPT's answer was the most trusted response by patients in two out of the four questions (Question #1 and #3).
Conclusion:
The study suggests that patients perceive ChatGPT as highly credible, particularly in terms of accuracy, authenticity, and believability compared with orthopaedic surgeons. These findings underscore the potential of AI to improve healthcare literacy and patient decision making. Further research is warranted to explore the nuances of patient trust and preferences in AI-generated healthcare content.
Level Of Evidence:
Level II, prospective, comparative study.

