Related Experiment Video
Updated: Jul 19, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Inherent Bias in Large Language Models: A Random Sampling Analysis
Noel F Ayoub1, Karthik Balakrishnan2, Marc S Ayoub3
1Division of Rhinology and Skull Base Surgery, Department of Otolaryngology--Head & Neck Surgery, Mass Eye and Ear/Harvard Medical School, Boston, MA.
Abstract:
There are mounting concerns regarding inherent bias, safety, and tendency toward misinformation of large language models (LLMs), which could have significant implications in health care. This study sought to determine whether generative artificial intelligence (AI)-based simulations of physicians making life-and-death decisions in a resource-scarce environment would demonstrate bias. Thirteen questions were developed that simulated physicians treating patients in resource-limited environments. Through a random sampling of simulated physicians using OpenAI's generative pretrained transformer (GPT-4), physicians were tasked with choosing only 1 patient to save owing to limited resources. This simulation was repeated 1000 times per question, representing 1000 unique physicians and patients each. Patients and physicians spanned a variety of demographic characteristics. All patients had similar a priori likelihood of surviving the acute illness. Overall, simulated physicians consistently demonstrated racial, gender, age, political affiliation, and sexual orientation bias in clinical decision-making. Across all demographic characteristics, physicians most frequently favored patients with similar demographic characteristics as themselves, with most pairwise comparisons showing statistical significance (P<.05). Nondescript physicians favored White, male, and young demographic characteristics. The male doctor gravitated toward the male, White, and young, whereas the female doctor typically preferred female, young, and White patients. In addition to saving patients with their own political affiliation, Democratic physicians favored Black and female patients, whereas Republicans preferred White and male demographic characteristics. Heterosexual and gay/lesbian physicians frequently saved patients of similar sexual orientation. Overall, publicly available chatbot LLMs demonstrate significant biases, which may negatively impact patient outcomes if used to support clinical care decisions without appropriate precautions.
Insights
Generative artificial intelligence (AI) simulations revealed significant physician bias in life-or-death decisions. Large language models (LLMs) favored patients similar to the simulated physician, impacting healthcare equity.
Area of Science:
- Medical Ethics and Artificial Intelligence
- Health Informatics and Bias Detection
Background:
- Growing concerns exist regarding inherent bias, safety, and misinformation potential of large language models (LLMs).
- These concerns have significant implications for the integration of AI in healthcare decision-making.
Purpose of the Study:
- To investigate whether generative artificial intelligence (AI)-based simulations of physicians exhibit bias in life-and-death decisions.
- To assess bias in resource-scarce clinical scenarios using AI simulations.
Main Methods:
- Developed 13 questions simulating physicians in resource-limited environments making critical treatment choices.
- Utilized OpenAI's GPT-4 to simulate 1000 unique physicians and patients per question, ensuring diverse demographics.
- Patients had similar a priori survival likelihoods; physicians chose one patient to save based on limited resources.
Main Results:
- Simulated physicians consistently demonstrated racial, gender, age, political affiliation, and sexual orientation bias.
- Physicians predominantly favored patients sharing their own demographic characteristics (P<.05).
- Specific biases observed included nondescript physicians favoring White, male, young patients; political affiliation influenced choices (Democrats favored Black/female; Republicans favored White/male).
Conclusions:
- Publicly available large language models exhibit significant biases in simulated clinical decision-making.
- These biases could negatively impact patient outcomes if AI tools are used in clinical support without safeguards.
- Urgent need for bias mitigation strategies in AI for healthcare applications.
More Related Videos
Related Concept Videos
Random Sampling Method
Random and Systematic Errors
Randomized Experiments
Simple randomization
Simple...
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
Systematic Error: Methodological and Sampling Errors
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Random Error

