在大型语言模型中测量性别和种族偏见:自动化简历评估的交叉证据
Jiafu An1, Difang Huang2, Chen Lin3
1Department of Real Estate and Construction, University of Hong Kong, Hong Kong SAR 999077, China.
PNAS nexus
|March 27, 2025
概括
大型语言模型 (LLM) 在求职者评估中显示性别和种族偏见. 虽然支持女性候选人,但LLM惩罚黑人男性候选人,可能会影响招聘多样性.
科学领域:
- 人工智能的人工智能
- 社会科学 社会科学 社会科学
- 计算机科学 计算机科学
背景情况:
- 人类的决策倾向于社会偏见,导致代表性不足的群体的经济结果不平等.
- 大型语言模型 (LLM) 的兴起表明,人工智能在决策中正在转向人工智能,这引发了关于人工智能对公平性的影响的问题.
研究的目的:
- 在评估入门级求职者时,调查常用LLM中的性别和种族偏见.
- 确定LLM偏差如何影响各种社会群体在高风险招聘场景中的分配结果.
主要方法:
- 士 (GPT-3.5 Turbo,GPT-4o,Gemini 1.5 Flash,Claude 3.5 Sonnet,Llama 3-70b) 被指示对大约361,000个简历进行评分.
- 简历中出现了随机的社会身份,以评估与性别和种族有关的偏见.
主要成果:
- 在LLMs中,具有相似资格的女性候选人获得更高的分数,而黑人男性候选人则获得更低的分数.
- 这些偏见可能导致相似候选人的招聘概率在1-3个百分点之间存在差异.
- 在不同的社会群体和职业类型中,偏见的方向和程度各不相同.
结论:
- 基于LLM的AI系统在招聘评估中表现出明显的性别和种族偏见.
- 解决和最大限度地减少这些人工智能偏见对于确保随着人工智能采用增长的公平结果至关重要.
- 需要进一步的研究来了解偏见的起源,并制定缓解策略.
相关概念视频
Stereotypes, Prejudice, and Discrimination
89.7K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
89.7K
Confirmation Biases
5.4K
The confirmation bias is the tendency to focus on information that confirms our existing beliefs and ignore information that is inconsistent with our expectations. For example, if you think that your professor is not very nice, you notice all of the instances of rude behavior exhibited by the professor while ignoring the countless pleasant interactions he is involved in on a daily basis. Have you ever fallen prey to the confirmation bias, either as the source or target of such bias?
5.4K
Stereotype Content Model
13.9K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
13.9K
Stereotype Threat and Self-fulfilling Prophecies
37.3K
When we hold a stereotype about a person, we have expectations that he or she will fulfill that stereotype. A self-fulfilling prophecy is an expectation held by a person that alters his or her behavior in a way that tends to make it true. When we hold stereotypes about a person, we tend to treat the person according to our expectations. This treatment can influence the person to act according to our stereotypic expectations, thus confirming our stereotypic beliefs. Research by Rosenthal and...
37.3K
Quantifying and Rejecting Outliers: The Grubbs Test
1.4K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.4K
Improving Translational Accuracy
2.5K
2.5K


