Related Experiment Video
Updated: Mar 27, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
When Artificial Intelligence Fails the Future: How Unchecked Artificial Intelligence Could Amplify Inequality in
Omar Allam1, Ismail Ajjawi1, Andrew Salib1
1From the Division of Plastic & Reconstructive Surgery, Department of Surgery, Yale School of Medicine, New Haven, CT.
Background:
Artificial intelligence is influencing healthcare but carries potential biases from historically trained data. This study was the first to examine racial and gender bias in large language model architecture applied to medical student career advising.
Methods:
Two hundred synthetic medical student profiles were created with systematically varied demographic and academic characteristics. Each profile was evaluated using a standardized advising prompt that generated the top 3 specialty recommendations. Specialties were assigned competitiveness scores based on National Resident Matching Program data. Univariate and multivariate analyses assessed associations between applicant characteristics and predicted competitiveness and surgical specialty recommendations. A compensatory scoring analysis estimated the additional USMLE Step 2 Clinical Knowledge points required for applicants from different demographic groups to achieve competitiveness equivalent to a White male reference profile.
Results:
Male applicants were rated more competitive than female applicants (P < 0.0001). Black and Hispanic applicants were deemed less competitive than White applicants (P = 0.0019 and P = 0.0108, respectively). Male applicants were more than 10 times more likely to receive surgical specialty recommendations than female applicants (odds ratio = 10.23, P < 0.0001). Black and Hispanic applicants were significantly less likely to receive surgical recommendations compared with White applicants. Compensatory scoring revealed that White female applicants needed 56 additional Step 2 Clinical Knowledge points to match perceived competitiveness, whereas Black and Hispanic female applicants required 95 and 88 additional points, respectively.
Conclusions:
Our findings demonstrated that, if left unchecked, large language models such as ChatGPT perpetuate racial and gender biases when advising medical students, amplifying historical inequities in medicine.
Related Concept Videos
Issues And Trends In Healthcare Delivery System
Cost Containment
Payment for healthcare services has historically promoted adoption of costly and often unnecessary or inefficient...
Documentation of Nursing Diagnosis
In some settings, data-driven computerized decision support systems are in place, allowing for more accurate nursing diagnoses. The database within one of these systems includes diagnostic labels defining characteristics, activities, and indicators for nursing. A nurse enters...
Current Trends in Nursing II
Non-equilibrium in the Cell