一个吉布斯后置框架,用于公平的集群
Abhisek Chakraborty1, Anirban Bhattacharya1, Debdeep Pati1
1Department of Statistics, Texas A&M University, College Station, TX 77843, USA.
Entropy (Basel, Switzerland)
|January 22, 2024
概括
本研究引入了公平聚类的新贝叶斯方法,通过提供不确定性量化和决策理论解释来解决现有方法的局限性. 该框架确保了算法公平性,而无需显著的计算成本.
科学领域:
- 机器学习 机器学习
- 算法公平性 算法公平性
- 集群集成是指集群集成.
背景情况:
- 算法公平性在机器学习中至关重要,在集群中使用"平衡"作为公平性标准.
- 现有的公平集群方法 (例如,k-means) 缺乏不确定性量化,并与模型错误规范作斗争.
- 基于混合模型的方法提供不确定性量化,但计算密集且脆弱.
研究的目的:
- 为公平的集群提出一种新的概率公式,使不确定性量化.
- 开发一个广义的贝叶斯公平集群框架与决策理论解释.
- 为公平的集群设计高效的计算算法.
主要方法:
- 开发了一个广义的贝叶斯框架,用于公平的集群.
- 设计了利用最佳运输和基于损失的集群技术的高效算法.
- 这种方法使得即使在轻微的模型错误规范下,也可以轻松量化不确定性.
主要成果:
- 拟议的贝叶斯框架为公平的集群提供了有效的不确定性量化.
- 开发的算法具有计算效率,克服了现有方法的局限性.
- 数字实验和现实世界数据证明了拟议方法的有效性.
结论:
- 新的贝叶斯公平集群框架为不确定性量化提供了一个强大的解决方案.
- 高效的算法使公平的集群更容易获得和可靠.
- 这项工作推进了机器学习应用中的算法公平领域.
相关概念视频
Cluster Sampling Method
11.9K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.9K
Friedman Two-way Analysis of Variance by Ranks
198
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
198
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Law of Independent Assortment
55.7K
While Mendel’s Law of Segregation states that the two alleles for one gene are separated into different gametes, a different question of how different genes are inherited remains. For example, is the gene for tall plants inherited with the gene for green peas? Mendel asked this question by experimenting with a dihybrid cross; a cross in which both parents are homozygous for two distinct traits resulting in an F1 generation that are heterozygous for both traits.
55.7K
In- and Out-Groups
39.0K
People all belong to a gender, race, age, and social economic group. These groups provide a powerful source of our identity and self-esteem (Tajfel & Turner, 1979) and serve as our in-groups. An in-group is a group that we identify with or see ourselves as belonging to.
39.0K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K


