Related Experiment Video
Updated: Jun 2, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Implicit Gender, Racial, and Ethnic Biases in Large Language Models: An Audit Study of Automated Psychiatric
Sachin R Pendse1,2, Mini Jain1, Neha Kumar1
1School of Interactive Computing, Georgia Institute of Technology, Atlanta, GA.
Objective:
To evaluate whether large language models (LLMs) exhibit implicit gender, racial, and ethnic biases when used to provide psychiatric diagnoses, and to understand how such biases may impact the accuracy and fairness of AI-assisted clinical decision support.
Patients And Methods:
We conducted a large-scale audit of 6 LLMs, including general-purpose and medical-specific models, using 97 Diagnostic and Statistical Manual of Mental Disorders 5 psychiatric training cases, conducted between October 1, 2023, and June 23, 2025. Cases were systematically altered to suggest different gender and racial/ethnic identities-across 39 demographic groups-by changing names, pronouns, and descriptors while keeping clinical symptoms constant. We assessed diagnostic accuracy, additional or missed diagnoses, and language of diagnostic reasoning. Of the LLMs tested, Generative Pretrained Transformer 4o emerged as the most accurate model and was selected for deeper analysis.
Results:
Although Generative Pretrained Transformer 4o accurately identified at least 1 correct diagnosis in 82.8% of cases (15,346 of 18,527 cases), it often added non-ground truth diagnoses (70.3% of cases, or 13,017 of 18,527 cases), suggesting a tendency to overdiagnose. Accuracy varied by gender, with higher performance for female patients and lower for nonbinary individuals. Although overall accuracy did not differ significantly by race/ethnicity, biased diagnostic patterns emerged. For example, cultural bereavement and antisocial behavior-were diagnosed exclusively in patients of color, and terms such as disruptive were used more frequently for Black men.
Conclusion:
Our findings found that LLMs reproduce and reinforce clinical biases even when symptoms are constant. Moreover, AI-based tools must be audited not only for accuracy but also for bias in both diagnoses and explanatory language, especially when used in high-stakes mental health contexts.
Related Concept Videos
Stereotype Content Model
Stereotypes, Prejudice, and Discrimination
Diagnostic and Statistical Manual of Mental Disorders (DSM)
Motivational Bias
Confirmation Biases
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
