Related Experiment Video
Updated: Jun 7, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
504
Benchmark suites instead of leaderboards for evaluating AI fairness
Angelina Wang1,2, Aaron Hertzmann3, Olga Russakovsky1
1Princeton University, Princeton, NJ, USA.
Patterns (New York, N.Y.)
|November 21, 2024
Summary
Leaderboards for artificial intelligence (AI) fairness are problematic. Researchers propose reformed benchmarks in curated suites to better understand AI fairness trade-offs and potential harms.
Area of Science:
- Artificial Intelligence
- Machine Learning Ethics
- AI Fairness
Background:
- Leaderboards and benchmarks are common tools for assessing artificial intelligence (AI) model fairness.
- Critics argue leaderboards incentivize optimizing for specific metrics, which is impossible due to varying application needs.
- Current critiques often entangle leaderboards and benchmarks, obscuring the value of benchmarks.
Purpose of the Study:
- To disentangle critiques of AI leaderboards and benchmarks.
- To propose a reformed approach to using benchmarks for AI fairness assessment.
- To advocate for the development of curated benchmark suites.
Main Methods:
- Analyzing the critiques of AI leaderboards and benchmarks.
- Conceptualizing the structure and purpose of benchmark suites.
- Outlining research directions for creating effective benchmark suites.
Main Results:
- Benchmarks, when separated from leaderboards, are valuable tools for understanding AI models.
- Curated benchmark suites can help researchers and practitioners identify a wide range of potential AI harms and fairness trade-offs.
- The proposed approach moves away from competitive leaderboards towards a more nuanced understanding of AI fairness.
Conclusions:
- Reformed benchmarks, organized into suites, offer a path to better monitor and improve AI fairness.
- Future research should focus on developing benchmark suites that address diverse usage, potential harms, and varied perspectives.
- Shifting focus from leaderboards to thoughtfully designed benchmark suites is crucial for advancing AI fairness.
Related Concept Videos
Bias
3.7K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
3.7K
Measures of Intelligence
6.4K
Psychologists measure intelligence by using standardized tests that produce a score known as the intelligence quotient or IQ. To understand IQ tests, it's important to recognize the key principles behind their construction: validity, reliability, and standardization.
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
Validity refers to how well a test measures what it claims to measure. An intelligence test should accurately assess intelligence rather than another characteristic, like anxiety. Criterion validity is one way to evaluate this;...
6.4K
Multiple Comparison Tests
3.9K
Multiple comparison test, abbreviated as MCT, is a post hoc analysis generally performed after comparing multiple samples with one or more tests. An MCT will help identify a significantly different sample among multiple samples or a factor among multiple factors.
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
It would be easy to compare two samples using a significance alpha level of 0.05. In other words, there is only one sample pair to be compared. However, it would be difficult to identify a significantly different sample if the number...
3.9K
Stereotype Content Model
14.0K
The Stereotype Content Model (SCM) was first proposed by Susan Fiske and her colleagues (Fiske, Cuddy, Glick & Xu, 2002; see also Fiske, 2012 and Fiske, 2017). The SCM specifies that when someone encounters a new group, they will stereotype them based on two metrics: warmth—or that group’s perceived intent, and how likely they are to provide help or inflict harm—and competence—or their ability to carry out that objective. Depending on the warmth-competence...
14.0K
Stereotypes, Prejudice, and Discrimination
90.0K
Humans are very diverse and although we share many similarities, we also have many differences. The social groups we belong to help form our identities (Tajfel, 1974). These differences may be difficult for some people to reconcile, which may lead to prejudice toward people who are different. Prejudice is a negative attitude and feeling toward an individual based solely on one’s membership in a particular social group (Allport, 1954; Brown, 2010). Prejudice is common against people who...
90.0K
Bonferroni Test
2.7K
The Bonferroni test is a statistical test named after Carlo Emilio Bonferroni, an Italian mathematician best known for Bonferroni inequalities. This statistical test is a type of multiple comparison test to determine which means are different than the rest. Bonferroni test can minimize the Type 1 error by reducing the significance level alpha, which otherwise increases with sample pairs.
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
The means of different samples are first paired in all possible combinations.
The null hypothesis of the...
2.7K

