基准套件而不是排名表来评估AI公平性
Angelina Wang1,2, Aaron Hertzmann3, Olga Russakovsky1
1Princeton University, Princeton, NJ, USA.
Patterns (New York, N.Y.)
|November 21, 2024
概括
人工智能 (AI) 公平性的排行榜存在问题. 研究人员在精选套件中提出了改革的基准,以更好地了解AI公平性权衡和潜在的危害.
科学领域:
- 人工智能的人工智能
- 机器学习伦理学 机器学习伦理学
- 人工智能 公平 公平
背景情况:
- 排行榜和基准是评估人工智能 (AI) 模式公平性的常见工具.
- 批评者认为,排行榜激励优化特定指标,这是不可能的,因为不同的应用需求.
- 当前的批评经常纠排行榜和基准,掩盖了基准的价值.
研究的目的:
- 解开对人工智能排行榜和基准的批评.
- 为AI公平性评估提出使用基准的改革方法.
- 倡导开发精心策划的基准套件.
主要方法:
- 分析AI排名表和基准的批评.
- 概念化基准套件的结构和目的.
- 概述研究方向,以创建有效的基准套件.
主要成果:
- 当与排名表分开时,基准是理解AI模型的有价值工具.
- 精心策划的基准套件可以帮助研究人员和从业人员识别各种潜在的AI伤害和公平性权衡.
- 拟议的方法远离了竞争的排行榜,转向了对人工智能公平性的更细致的理解.
结论:
- 改革后的基准,组织成套件,提供了一条更好地监测和改善AI公平性的途径.
- 未来的研究应该专注于开发针对不同用途,潜在危害和不同观点的基准套件.
- 将焦点从排行榜转移到经过深思熟虑设计的基准套件对于推进AI公平性至关重要.
相关概念视频
Bias
3.7K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.7K
Measures of Intelligence
6.4K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
6.4K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Stereotypes, Prejudice, and Discrimination
90.0K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
90.0K
Bonferroni Test
2.7K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.7K


