在医疗报告生成中揭露和量化大型语言模型的种族偏见
Yifan Yang1,2, Xiaoyu Liu2, Qiao Jin1
1National Institutes of Health (NIH), National Library of Medicine (NLM), National Center for Biotechnology Information (NCBI), Bethesda, MD, 20894, USA.
Communications medicine
|September 10, 2024
概括
在医疗保健中使用的大型语言模型 (LLM) 显示出偏见,预测白人人口的成本更高,住院时间更长. 解决这些人工智能偏见对于公平的患者护理至关重要.
科学领域:
- 人工智能在医学中的应用
- 自然语言处理自然语言处理.
- 医疗保健的不平等 医疗保健的不平等
背景情况:
- 像GPT-3.5-turbo和GPT-4这样的大型语言模型 (LLM) 在医疗保健中提供了潜在的好处.
- 然而,LLM可能会从训练数据中继承并延续偏见,影响医疗应用.
- 这些偏见在医疗保健中的确切程度和影响仍未得到充分研究.
研究的目的:
- 在医疗背景下调查大型语言模型 (LLM) 中的偏差.
- 评估LLM反映和放大现有的医疗保健差异的潜力.
主要方法:
- 利用LLM生成使用真实患者病例的住院,成本和死亡率的预测.
- 对LLM生成的响应进行手动检查,以识别和分析偏见.
- 在患者背景生成,疾病组关联和治疗建议中评估偏差表现.
主要成果:
- 法律法规显示偏见,预测白人人口的成本更高,住院时间更长.
- 模型在具有挑战性的医疗场景中显示出乐观的生存率预测.
- 偏见反映了现实世界的医疗保健差异,影响了患者的背景生成和治疗建议.
结论:
- 这些发现突显了与医疗保健应用相关的LLM中的重大偏差.
- 迫切需要研究临界医疗使用LLM的偏差缓解策略.
- 对所有患者群体来说,确保人工智能驱动的医疗保健的公平性和准确性至关重要.
相关概念视频
Bias in Epidemiological Studies
175
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
175
Bias
3.9K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.9K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Language and Cognition
336
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
336


