Related Experiment Video
Updated: Jan 18, 2026

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
1.3K
Race, Ethnicity and Their Implication on Bias in Large Language Models
Shiyue Hu1,2, Ruizhe Li3, Yanjun Gao1
1University of Colorado Anschutz.
Medrxiv : the Preprint Server for Health Sciences
|January 16, 2026
Summary
This study reveals how large language models (LLMs) process race and ethnicity, finding that bias mitigation through neuron intervention shows limited success, indicating deeper representational issues in AI models.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Computational Linguistics
Background:
- Large language models (LLMs) are increasingly used in sensitive areas like healthcare.
- Existing research highlights outcome disparities but lacks insight into internal LLM mechanisms.
- Understanding how LLMs represent demographic attributes is crucial for mitigating bias.
Purpose of the Study:
- To investigate the internal mechanisms by which large language models represent and operationalize race and ethnicity.
- To analyze the distribution and function of demographic information within LLM architectures.
- To assess the impact of targeted interventions on bias reduction.
Main Methods:
- Analysis of three open-source LLMs using two public datasets (toxicity generation, clinical narrative understanding).
- Application of a reproducible interpretability pipeline combining probing, neuron-level attribution, and targeted intervention.
- Examination of how demographic cues influence model behavior and internal representations.
Main Results:
- Demographic information is distributed across internal model units with significant cross-model variability.
- Some units exhibit stereotype-associated learning from pretraining data.
- Identical demographic cues can trigger different model behaviors, and interventions yield incomplete bias reduction.
Conclusions:
- LLM internal representations of race and ethnicity are complex and vary across models.
- Bias mitigation strategies targeting specific neurons show limited effectiveness, suggesting deeper representational challenges.
- Further research is needed for more systematic approaches to address bias in LLMs.
Related Concept Videos
Stereotypes, Prejudice, and Discrimination
95.0K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
95.0K
Language and Cognition
716
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
716
Bias
7.2K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
7.2K
Surveys
16.6K
Often, psychologists develop surveys as a means of gathering data. Surveys are lists of questions to be answered by research participants, and can be delivered as paper-and-pencil questionnaires, administered electronically, or conducted verbally. Generally, the survey itself can be completed in a short time, and the ease of administering a survey makes it easy to collect data from a large number of people.
16.6K
Motivational Bias
313
Cognitive bias results from limitations in thinking and information processing, leading to systematic errors in judgment. Conversely, motivational bias stems from personal desires or emotions, causing distortions in perception to align with self-interest. Motivational bias influences how individuals perceive and attribute causes to events, often shaped by personal needs, goals, and self-esteem preservation. This bias can distort judgment, leading to inaccurate assessments of success, failure,...
313
Stereotype Content Model
15.3K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
15.3K

