Related Experiment Video
Updated: Jan 7, 2026

Evidence-based Knowledge Synthesis and Hypothesis Validation: Navigating Biomedical Knowledge Bases via Explainable AI and Agentic Systems
Published on: June 13, 2025
Uncovering bias and variability in how large language models attribute cardiovascular risk
Justine Tin Nok Chan1, Ray Kin Kwek2
1School of Clinical Medicine, University of Cambridge, Cambridge, United Kingdom.
Insights
Large language models (LLMs) show bias in cardiovascular risk assessment, attributing higher risk to men and Black patients. Their decisions varied with comorbidities, highlighting the need for careful evaluation to prevent health inequities.
Area of Science:
- Medical Artificial Intelligence
- Clinical Decision Support Systems
- Health Equity Research
Background:
- Large language models (LLMs) are increasingly integrated into medical applications.
- However, their performance in attributing cardiovascular risk, particularly concerning demographic and clinical factors, is not well understood.
- This gap necessitates an examination of LLM decision-making processes in this critical area.
Purpose of the Study:
- To investigate how a specific LLM (ChatGPT 4.0 mini) assigns relative cardiovascular risk across various demographic and clinical domains.
- To assess the LLM's consistency and potential biases in risk attribution.
- To explore the impact of specific clinical factors (e.g., BMI, diabetes, depression, smoking, hyperlipidemia) on the LLM's risk assessments.
Main Methods:
- A structured set of prompts was designed covering six domains: general cardiovascular risk, BMI, diabetes, depression, smoking, and hyperlipidemia.
- Prompts were submitted in triplicate to ChatGPT 4.0 mini.
- Neutral prompts assessed baseline risk attribution, while comparative prompts evaluated changes in risk assignment when specific domains were included.
Main Results:
- The LLM generally attributed higher cardiovascular risk to men than women and to Black patients compared to white patients across neutral prompts.
- In comparative analyses, sex-based risk attributions shifted in two of six domains (e.g., with depression, risk was equal; with smoking, males were higher risk).
- Race-based risk attributions remained consistent, with the LLM consistently identifying Black patients as higher risk. High agreement across repeated runs (ICC=0.949) was observed.
Conclusions:
- The LLM demonstrated both bias and variability in its cardiovascular risk attributions across different domains.
- While sex-based risk perceptions could be influenced by comorbidities, race-based perceptions were notably stable.
- These findings underscore the critical need for rigorous evaluation of LLMs in clinical settings to mitigate the risk of perpetuating existing health inequities.
Abstract:
Large language models (LLMs) are used increasingly in medicine, but their decision-making in cardiovascular risk attribution remains underexplored. This pilot study examined how an LLM apportioned relative cardiovascular risk across different demographic and clinical domains. A structured prompt set across six domains was developed, across general cardiovascular risk, body mass index (BMI), diabetes, depression, smoking, and hyperlipidaemia, and submitted in triplicate to ChatGPT 4.0 mini. For each domain, a neutral prompt assessed the LLM's risk attribution, while paired comparative prompts examined whether including the domain changed the LLM's decision of the higher-risk demographic group. The LLM attributed higher cardiovascular risk to men than women, and to Black rather than white patients, across most neutral prompts. In comparative prompts, the LLM's decision between sex changed in two of six domains: when depression was included, risk attribution was equal between men and women. It changed from females being at higher risk than males in scenarios without smoking, but changed to males being at higher risk than females when smoking was present. In contrast, race-based decisions of relative risk were stable across domains, as the LLM consistently judged Black patients to be higher-risk. Agreement across repeated runs was strong (ICC of 0.949, 95% CI: 0.819-0.992, p = <0.001). The LLM exhibited bias and variability across cardiovascular risk domains. Although decisions between males/females sometimes changed when comorbidities were included, race-based decisions remained the same. This pilot study suggests careful evaluation of LLM clinical decision-making is needed, to avoid reinforcing inequities.
More Related Videos
Related Concept Videos
Bias in Epidemiological Studies
Factors Influencing Heart Rate
Let us explore the significant factors affecting heart rate, including age, body temperature, posture, acute pain, chemical influences,...
Confounding in Epidemiological Studies
Variability: Analysis
The range is a simple measure of variability, indicating the difference between the highest and...
Language and Cognition
Coronary Artery Disease I: Introduction

