Related Experiment Video
Updated: Jul 12, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Bias and Inaccuracy in AI Chatbot Ophthalmologist Recommendations
Michael C Oca1, Leo Meller1, Katherine Wilson1
1Orthopedic Surgery, Shiley Eye Institute, University of California (UC) San Diego Health, La Jolla, USA.
AI chatbots like ChatGPT, Bing Chat, and Google Bard show significant bias and inaccuracy when recommending ophthalmologists. These tools often fail to provide accurate physician referrals and exhibit gender bias, particularly against female doctors.
Area of Science:
- Artificial Intelligence in Healthcare
- Medical Informatics
- Ophthalmology Practice Management
Background:
- AI chatbots are increasingly used for information retrieval, including healthcare-related queries.
- The accuracy and potential biases of AI recommendations in specialized medical fields require thorough evaluation.
- Understanding AI performance in physician referrals is crucial for patient safety and equitable access to care.
Purpose of the Study:
- To assess the accuracy and identify biases in ophthalmologist recommendations from ChatGPT 3.5, Bing Chat, and Google Bard.
- To analyze chatbot performance across the 20 most populous U.S. cities.
- To compare AI-generated recommendations against national averages for physician demographics and practice types.
Main Methods:
- Generated 80 recommendations per chatbot using the prompt 'Find me four good ophthalmologists in (city)'.
- Collected physician characteristics: specialty, location, gender, practice type, and fellowship.
- Utilized one-proportion z-tests and Pearson's chi-squared tests to compare chatbot recommendations with national data (e.g., proportion of female ophthalmologists, academic practice prevalence).
Main Results:
- Bing Chat (1.61%) and Google Bard (8.0%) significantly underrepresented female ophthalmologists compared to the national average (27.2%).
- All chatbots demonstrated high inaccuracy rates: ChatGPT (73.8%), Bing Chat (67.5%), and Bard (62.5%).
- Chatbots disproportionately recommended ophthalmologists in academic medicine or combined practices, exceeding the national academic average (17%).
Conclusions:
- AI chatbots exhibit significant bias and inaccuracy in recommending ophthalmologists, often suggesting physicians outside the specialty or desired location.
- Bing Chat and Google Bard displayed a bias against recommending female ophthalmologists.
- All evaluated chatbots showed a preference for recommending ophthalmologists in academic settings.
More Related Videos
05:47Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
05:49Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders
Published on: November 1, 2024
Related Concept Videos
Open Angle Glaucoma: Treatment
Drugs such as carbonic anhydrase inhibitors, α2- and...
Angle Closure Glaucoma: Treatment
Glaucoma: Overview
The Availability Heuristic
Non-equilibrium in the Cell
Accuracy and Errors in Hypothesis Testing
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...