Related Experiment Video
Updated: Jul 2, 2026

05:57
A Teleoperated Robotic System-Assisted Percutaneous Transiliac-Transsacral Screw Fixation Technique
Published on: January 6, 2023
Prompt engineering experiment on ChatGPT's ability to recommend orthopedic surgeons
Michael Robert Haupt1, Daniel Massillon2, Luning Yang3
1Global Health Program, Department of Anthropology, University of California, San Diego, CA USA; Global Health Policy & Data Institute, San Diego, CA USA.
International Journal of Medical Informatics
|June 30, 2026
Summary
Large language models like ChatGPT often fail to provide valid doctor recommendations. Their responses can be biased by patient characteristics, impacting healthcare access and equity.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Health Services Research
Background:
- Patients increasingly use online resources and AI tools for healthcare provider selection.
- Large language models (LLMs) are being explored for their potential to assist in finding medical practitioners.
- Challenges exist in ensuring AI-generated recommendations are accurate, unbiased, and equitable.
Purpose of the Study:
- To evaluate the accuracy and reliability of ChatGPT in recommending orthopedic surgeons.
- To investigate the impact of patient characteristics on ChatGPT's recommendation responses.
- To identify potential biases in LLM-driven healthcare provider recommendations.
Main Methods:
- Conducted 40,500 queries to ChatGPT for orthopedic surgeon recommendations.
- Varied patient persona characteristics (age, race, income, insurance, location) in prompts.
- Analyzed response rates, validity of recommendations, and demographic characteristics of recommended surgeons.
Main Results:
- ChatGPT was unable to provide recommendations for 52.8% of queries.
- High income and health insurance status significantly increased the likelihood of receiving a recommendation.
- Recommendations varied based on patient race and geographic location (NYC/Chicago more likely than Phoenix/Houston).
- Only 44.6% of recommended surgeons were valid; 97.1% of valid recommendations were male and 81.4% were White.
Conclusions:
- ChatGPT does not reliably provide valid orthopedic surgeon recommendations.
- LLM recommendations are influenced by patient demographics and socioeconomic factors, potentially exacerbating healthcare disparities.
- Further research is needed to address biases and improve the safety and equity of AI in healthcare decision-making.
