人工智能生成的模拟患者队列中的人口偏差:与人口普查基准对比的比较分析
Miriam Veenhuizen1, Andrew O'Malley2
1School of Medicine, University of St Andrews, North Haugh, St Andrews, KY16 9TF, UK.
Advances in simulation (London, England)
|November 19, 2025
概括
生成型人工智能模型产生人口统计学上狭窄的合成患者队列,缺乏年龄,性别和种族的多样性. 在模拟患者中的这种偏见可能会对医学教育和公平的临床实践产生负面影响.
科学领域:
- 医疗教育 技术 技术 医学教育
- 医疗保健中的人工智能
- 计算人口统计学 计算人口统计学
背景情况:
- 生成型人工智能 (AI) 模型为本科医学教育提供低成本的模拟患者队列.
- 这些人工智能模型的教育实用性取决于它们反映现实世界人口多样性的能力.
- 本研究评估了由两种大型语言模型 (LLM) 产生的合成患者资料的人口代表性.
研究的目的:
- 评估通常部署的LLM是否产生反映当前年龄,性别和种族组成的合成英国患者个人资料.
- 在医学培训中使用的人工智能生成的患者队列中识别人口偏见.
- 在人工智能驱动的教育工具中,为改进人口统计准确性制定战略提供信息.
主要方法:
- 两个LLM (GPT-3.5-turbo-0125和GPT-4-mini-2024-07-18) 产生了250个英国患者个人资料,没有人口统计指导.
- 年龄是直接获得的;性别和种族是从名字中推断出来的,使用一个验证的分类器.
- 观察到的人口统计频率与英格兰和威尔士2021年人口普查数据进行了比较,使用千平方适应性测试.
主要成果:
- 两种人工智能模型都产生了与人口普查基准显著分离的合成患者队列,涉及所有人口统计变量 (p < 0.0001).
- 年龄分布是倾斜的,缺乏非常年轻和年长的个体,在某些中年人群中占比过高;两种模型都没有产生25岁以下的患者,GPT-3.5的最老是47,GPT-4-mini的56.
- 性别比例严重倾向于男性 (64.7%和92.8%).
- 观察到的名称多样性有限,种族资料不平衡,一些群体的代表过多,而另一些群体则完全不存在.
结论:
- 评估的LLM在默认状态下产生了缺乏年轻人,老年人,女性和大多数少数民族代表性的合成患者池.
- 人口狭窄的人工智能输出有可能使偏见的临床期望正常化,并破坏公平医疗实践的培训.
- 在将生成系统集成到医学课程之前,对人工智能模型行为进行基线审计对于评估提示工程和数据策划策略至关重要.
更多相关视频
相关概念视频
Bias in Epidemiological Studies
1.2K
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
1.2K
Bias
7.2K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
7.2K
Stereotypes, Prejudice, and Discrimination
94.8K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
94.8K
Analysis of Population Pharmacokinetic Data
659
Analysis of population pharmacokinetic data involves studying the behavior of drugs within diverse populations to understand their pharmacokinetic parameters. Traditional pharmacokinetic methods typically involve collecting samples from a few individuals and estimating these parameters. While these methods are commonly used, they have limitations in capturing the variability in drug response among individuals or heterogeneous populations. Population pharmacokinetics is employed to address these...
659
Confounding in Epidemiological Studies
556
Confounding in statistical epidemiology represents a pivotal challenge, referring to the distortion in the perceived relationship between an exposure and an outcome due to the presence of a third variable, known as a confounder. This variable is associated with both the exposure and the outcome but is not a direct link in their causal chain. Its presence can lead to erroneous interpretations of the exposure's effect, either exaggerating or underestimating the true association. This...
556
Regression Toward the Mean
6.8K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.8K


