Related Experiment Video
Updated: Jan 6, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Invisible Bias in GPT-4o-mini: Detecting Disparities in AI-Generated Patient Messaging.
Christopher Reagen1, Katelyn M Brown2, Thomas Hocker3
1School of Medicine, University of Missouri-Kansas City, 2411 Holmes St, Kansas City, MO, 64108, USA. cjr8zx@umsystem.edu.
Artificial intelligence (AI), specifically large language models (LLMs), can enhance patient care but may contain hidden biases. This study found patient sex and religious preference significantly impacted AI-generated communication scores, highlighting potential disparities.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Natural Language Processing
Background:
- Large language models (LLMs) show increasing promise for improving patient care.
- However, biases within AI training data can lead to harmful disparities in medical communications.
- Evaluating these biases is crucial for safe AI implementation in medicine.
Purpose of the Study:
- To develop and propose a novel method for evaluating bias in LLMs.
- To assess AI-generated patient communications for disparities using synthetic patient data.
- To identify specific patient characteristics that may influence AI communication quality.
Main Methods:
- Utilized GPT-4o-mini to generate patient communications from synthetic medical records.
- Systematically created diverse synthetic patient data, including demographic and historical information.
- Employed GPT-4o-mini to score AI-generated communications on empathy, accuracy, clarity, professionalism, respect, and encouragement.
- Analyzed scoring disparities linked to patient attributes to detect potential biases.
Main Results:
- Patient sex and religious preference demonstrated a statistically significant impact on the quality scores of AI-generated communications.
- Identified specific patient history components that correlated with disparities in AI communication.
- The proposed method successfully detected potential biases in LLM-generated patient interactions.
Conclusions:
- The study presents a novel methodology for bias detection in LLMs using synthetic data and AI-based scoring.
- Findings indicate that LLMs like GPT-4o-mini may exhibit biases related to patient demographics and personal history.
- Further research is recommended to refine evaluation criteria and assess a broader range of LLMs for clinical application.
Related Concept Videos
Stereotypes, Prejudice, and Discrimination
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Halo Effect
Barriers to Effective Communication II
Cultural barriers:
Differences in values, beliefs, religion, knowledge, and tradition can significantly impact communication. Awareness of nonverbal cues is critical, especially when conversing with a patient from a different culture. What appears appropriate in one culture may be inappropriate in another.
Semantic barriers:
As a result of their tendency to use...
Confirmation Biases
Impression Management Techniques IV: Altercasting

