Assessing Accuracy of Chat Generative Pre-Trained Transformer's Responses to Common Patient Questions Regarding
Niklaus P Zeller1, Ayush D Shah1, Ann E Van Heest1,2
1University of Minnesota Medical School, Minneapolis, MN.
Journal of Hand Surgery Global Online
|June 16, 2025
Summary
Chat Generative Pre-Trained Transformer (ChatGPT) 4.0 offers reliable answers to patient questions on congenital upper limb differences (CULDs). However, responses vary in depth and rarely direct patients to hand surgeons, necessitating cautious use for CULDs information.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Pediatric Orthopedics
Background:
- Congenital upper limb differences (CULDs) present unique challenges for patients and families.
- Accurate and accessible information is crucial for managing CULDs and treatment options.
- Large language models (LLMs) like ChatGPT offer potential for patient education but require validation.
Purpose of the Study:
- To evaluate the accuracy and reliability of ChatGPT 4.0 in answering frequently asked questions (FAQs) about CULDs.
- To assess the quality of ChatGPT 4.0's responses regarding CULDs and their treatment.
Main Methods:
- Two pediatric hand surgeons identified common FAQs from parents regarding CULDs.
- Sixteen FAQs covering syndactyly, polydactyly, radial longitudinal deficiency, thumb hypoplasia, and general CULDs were input into ChatGPT-4.0.
- Responses were graded on a 1-4 scale for quality, with independent chats used to prevent bias.
Main Results:
- ChatGPT 4.0 provided mostly reliable, evidence-based responses to CULDs FAQs, with 51% requiring no clarification.
- Significant variability in response depth was noted, and only 8% of responses were rated unsatisfactory.
- ChatGPT frequently recommended consulting healthcare providers (81 times) but rarely referred patients to hand surgeons (9 times).
Conclusions:
- ChatGPT 4.0 can offer evidence-based answers for a majority of CULDs FAQs.
- Variability in response comprehensiveness and a low referral rate to hand surgeons indicate a need for caution.
- LLMs should be used cautiously as patient education tools for CULDs, as information may not be consistently individualized or comprehensive.


