医疗决策中的社会人口学偏见通过大型语言模型
Mahmud Omar1,2,3, Shelly Soffer4, Reem Agbareia5
1The Windreich Department of Artificial Intelligence and Human Health, Icahn School of Medicine at Mount Sinai and the Mount Sinai Health System, New York, NY, USA. mahmudomar70@gmail.com.
医疗保健中的大型语言模型 (LLM) 可能会根据患者的社会人口统计数据推偏见的护理. 研究表明,LLM引导某些群体接受紧急护理或高级成像,反映了模型偏见,而不是临床需要.
科学领域:
- 人工智能在医学中的应用
- 健康 公平 研究 健康 公平 研究
- 临床决策支持系统 临床决策支持系统
背景情况:
- 大型语言模型 (LLM) 在医疗保健方面具有潜力,但对偏见的临床建议提出了担忧.
- 社会人口统计学特征可能会过度影响LLM产生的医疗建议,可能会加剧健康差异.
研究的目的:
- 评估大型语言模型 (LLM) 中基于患者社会人口统计因素的临床护理建议的偏见.
- 确定LLM输出是否反映了急诊部病例管理中的医学上不合理的差异.
主要方法:
- 分析了来自1000个急诊室病例 (真实和合成) 的170多万个LLM输出.
- 病例呈现了32种变化,改变了社会人口统计学标签,同时保持了临床细节的不变.
- 将LLM建议与来自医生的基线和模型特定的控制案例进行比较.
主要成果:
- 被标记为具有特定社会人口统计特征的病例 (例如,黑人,无家可归者,LGBTQIA+) 被不成比例地推用于紧急护理,侵入性干预或心理健康评估.
- 在某些LGBTQIA+子组中,LLMs建议对某些LGBTQIA+子组进行心理健康评估的频率是临床指标的6~7倍.
- 高收入的标记病例接受了更先进的成像建议,而低收入/中等收入的病例接受的更少,即使经过统计纠正.
结论:
- 观察到的LLM推差异没有得到临床推理的支持,这表明固有的模型偏差.
- 这种对LLM成果的偏见可能会导致健康差异,破坏公平的患者护理.
- 迫切需要在医疗保健应用的LLM中进行强有力的偏见评估和缓解策略.
更多相关视频
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
相关概念视频
Language and Cognition
Bias in Epidemiological Studies
Stereotype Content Model
Stereotypes, Prejudice, and Discrimination
Barriers to Effective Communication II
Cultural barriers:
Differences in values, beliefs, religion, knowledge, and tradition can significantly impact communication. Awareness of nonverbal cues is critical, especially when conversing with a patient from a different culture. What appears appropriate in one culture may be inappropriate in another.
Semantic barriers:
As a result of their tendency to use...
Bias
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
