Related Experiment Videos
Artificial Intelligence Can Direct Patients Toward a Complaint-specific Musculoskeletal Provider
Ethan C Gazan1, Colin M Emrich, Alexander J Baur
1From the Liberty University College of Osteopathic Medicine (Gazan, Emrich, Dr. Baur, Dr. Landy), Lynchburg, VA; the Department of Orthopaedic Surgery (Dr. Lum), University of California Davis, Sacramento, CA; the Connecticut Orthopaedics (Dr. Bernstein), Trumbull, CT, and the OrthoVirginia (Dr. Landy), Lynchburg, VA.
Summary
Large language models (LLMs) show potential for directing patients to musculoskeletal providers, but accuracy varies. ChatGPT demonstrated the highest appropriateness, while phone number accuracy differed significantly among AI models.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Digital Health
Background:
- Patients with musculoskeletal issues frequently use online resources to find healthcare providers.
- Large language models (LLMs) offer a novel approach to assist patients in identifying appropriate specialists.
- This study assesses the efficacy of LLMs in recommending musculoskeletal providers based on user queries.
Purpose of the Study:
- To evaluate the accuracy and appropriateness of provider recommendations made by three leading LLMs for musculoskeletal conditions.
- To compare the performance of ChatGPT, DeepSeek, and Gemini in identifying suitable healthcare providers.
- To assess the accuracy of contact information provided by LLMs for recommended physicians.
Main Methods:
- Standardized musculoskeletal patient queries were input into three LLMs (ChatGPT, DeepSeek, Gemini) for two distinct US cities.
- Provider recommendations were evaluated for appropriateness based on current practice and specialty.
- The accuracy of listed phone numbers was verified.
- Descriptive statistics and Fisher exact tests were employed for data analysis.
Main Results:
- ChatGPT achieved 100% appropriateness in provider recommendations, significantly outperforming Gemini (43%) and DeepSeek (40%) (P < 0.001).
- Inappropriate recommendations included real physicians in unrelated specialties (72%) and AI hallucinations (28%, exclusively from DeepSeek).
- Gemini exhibited the highest phone number accuracy (83%), followed by ChatGPT (67%) and DeepSeek (20%).
Conclusions:
- LLMs demonstrate promise in guiding patients to local, specialized musculoskeletal care, though contact information accuracy requires improvement.
- Healthcare providers should be cognizant of AI's role in patient navigation and ensure their practice information is accessible to these systems.
- Continued evolution of LLMs necessitates attention to data accuracy for reliable healthcare provider recommendations.