交叉护理:评估预培训数据对语言模型偏差的医疗保健影响
Shan Chen1,2,3, Jack Gallifant4, Mingye Gao4
1Harvard.
概括
大型语言模型 (LLM) 在人口统计学中的疾病流行率代表性中显示出显著的偏差. 交叉护理基准揭示了这些不准确性,冒着偏见的医疗AI应用的风险.
科学领域:
- 人工智能的人工智能
- 医疗信息学 医疗信息学
- 计算语言学 计算语言学
背景情况:
- 大型语言模型 (LLM) 对自然语言处理至关重要,但通常会从训练数据中继承偏差.
- 由于数据限制,现有的LLM可能会在医疗保健等敏感领域延续不准确性.
研究的目的:
- 介绍Cross-Care,这是一个新的基准框架,用于评估LLM偏见和现实世界的知识.
- 具体评估美国不同人口群体中疾病流行率的LLM代表性.
- 量化LLM输出与实际疾病流行数据之间的差异.
主要方法:
- 在LLM中使用预培训机构 (例如ThePile) 系统评估人口偏差.
- 将LLM产生的疾病患病率与美国人口统计数据的现实世界流行病学数据进行比较.
- 对调整方法在缓解表示不一致性方面的有效性进行分析.
主要成果:
- 在LLM疾病流行率代表和人口统计学子组的实际率之间发现了实质性的错位.
- 在医疗应用的LLM中显示偏差传播的风险.
- 在调整方法后,观察到跨语言疾病流行率一致性的最小改善.
结论:
- 法律法规表现出显著的偏见,并且在代表人口统计学范围内的疾病流行率方面缺乏现实世界的依据.
- 目前的对齐技术为这些特定偏差提供了有限的解决方案.
- 交叉护理框架为评估和减轻医疗AI偏见提供了一个关键工具.
更多相关视频
相关概念视频
Language and Cognition
287
Language serves as a bridge between ideas and communication, influencing how individuals perceive and interact with the world. Psychologists have long debated whether language shapes thought or vice versa. This discussion gained grip with Edward Sapir and Benjamin Lee Whorf in the 1940s, who proposed that language determines thought, a concept known as linguistic determinism. They suggested that the vocabulary and structure of a language influence how its speakers think and perceive reality.
287
Improving Translational Accuracy
2.5K
2.5K
Language Development
280
Children master language quickly and with relative ease, supported by both biological predisposition and reinforcement. B. F. Skinner (1957) proposed that language is learned through reinforcement, while Noam Chomsky (1965) argued that language acquisition mechanisms are biologically determined.
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
The critical period for language acquisition suggests that the ability to acquire language is at its peak early in life. As people age, this proficiency decreases. Language development begins very...
280
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Bias
3.7K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.7K
Data Validation
4.8K
Data validation is an essential part of a comprehensive assessment. Validation is confirming or verifying and opening the door to gathering more assessment data as it clarifies vague or unclear data. The process of checking and verifying the collected information is called data validation. The primary purpose of data validation is to ensure data is as free from error, bias, and misinterpretation as possible.
Nursing assessment guides are generally based on holistic models rather than medical...
Nursing assessment guides are generally based on holistic models rather than medical...
4.8K


