Related Experiment Video
Updated: Sep 12, 2025

05:47
Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
578
Evaluating gender bias in large language models in long-term care
1Care Policy and Evaluation Centre, LSE, London, WC2A 2AE, UK. s.w.rickman@lse.ac.uk.
BMC Medical Informatics and Decision Making
|August 10, 2025
Summary
State-of-the-art large language models (LLMs) show varying gender bias in summarizing long-term care records. Google Gemma exhibited significant bias, downplaying women's needs, unlike Llama 3.
Area of Science:
- Artificial Intelligence
- Natural Language Processing
- Healthcare Informatics
Background:
- Large language models (LLMs) are increasingly used to automate administrative tasks in long-term care, such as summarizing patient records.
- However, LLMs can perpetuate biases present in their training data, potentially impacting healthcare equity.
Purpose of the Study:
- To evaluate gender bias in summaries of long-term care records generated by recent open-source LLMs.
- Specifically comparing Meta's Llama 3 and Google Gemma against older benchmark models.
Main Methods:
- Generated gender-swapped versions of 617 long-term care records.
- Produced summaries using Llama 3, Gemma, T5, and BART.
- Quantified counterfactual gender bias using sentiment analysis, word frequency, and thematic patterns.
Main Results:
- Benchmark models showed some gender-based variations.
- Llama 3 demonstrated no significant gender-based differences in summaries.
- Google Gemma exhibited the most pronounced gender bias, with male summaries focusing more on health issues and women's needs being downplayed.
Conclusions:
- Gender bias in LLM summaries, particularly downplaying women's health issues, could lead to disparities in care service allocation.
- While LLMs offer administrative benefits, their varying performance necessitates rigorous bias evaluation.
- The study provides a practical framework for quantitatively assessing gender bias in LLMs.
More Related Videos
Related Concept Videos
Bias in Epidemiological Studies
676
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
676
Stereotype Content Model
14.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.9K
Stereotypes, Prejudice, and Discrimination
91.5K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
91.5K
Bias
4.9K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
4.9K
Language and Cognition
441
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
441
Lateralization
481
Brain lateralization refers to the division of mental processes and functions between the two hemispheres of the brain, a phenomenon that optimizes neural efficiency and underpins complex abilities in humans. This specialization allows each hemisphere to perform tasks where it has a comparative advantage, facilitating more refined cognitive capabilities across different domains.
481

