在帕雷托边界对机器学习的公平数据表示
1Department of Mathematics, University of California Davis, Davis, CA 95616-5270, USA.
概括
本研究介绍了一种新的预处理算法,用于公平的机器学习数据表示. 它优化了预测错误和统计差异之间的权衡,提高了AI决策中的公平性.
科学领域:
- 机器学习 机器学习
- 人工智能的人工智能
- 数据科学数据科学数据科学
背景情况:
- 机器学习 (ML) 越来越多地用于决策,需要公平的数据处理.
- 确保ML模型中的公平性至关重要,以防止有偏见的结果.
研究的目的:
- 提出一种预处理算法,以便在ML中公平地表示数据.
- 建立一种优化预测错误和统计差异之间的帕雷托边界的方法.
主要方法:
- 使用最佳的亲系传输来对瓦瑟斯坦重心的表征.
- 采用预处理数据变形以提供公平的表示.
- 分析瓦瑟斯坦的地质测量图来描述帕雷托边界.
主要成果:
- 预处理算法与各种监督学习方法和看不见的数据兼容.
- 公平代表性限制了从剩余数据中推断敏感信息的可能性.
- 最佳的亲属图表表现出计算效率,即使是高维数据.
结论:
- 建议的预处理方法有效地实现了公平的数据表示.
- 该方法提供了一种计算高效的方式来平衡预测准确性和公平性.
- 这项工作有助于开发更公平的AI系统.
相关概念视频
Pareto Chart
6.7K
A Pareto chart is a bar graph or a combination of both line and bar graphs. The bar lengths represent the individual values or the frequency, while the lines represent the cumulative total values. In this chart, the longest bars are arranged on the left and the shortest bars on the right, which makes it easier to read and interpret the data. It can also be called a Pareto diagram or Pareto analysis.
The Pareto chart is named after the Italian economist Vilfredo Pareto, who described the Pareto...
The Pareto chart is named after the Italian economist Vilfredo Pareto, who described the Pareto...
6.7K
Outliers and Influential Points
4.0K
An outlier is an observation of data that does not fit the rest of the data. It is sometimes called an extreme value. When you graph an outlier, it will appear not to fit the pattern of the graph. Some outliers are due to mistakes (for example, writing down 50 instead of 500), while others may indicate that something unusual is happening. Outliers are present far from the least squares line in the vertical direction. They have large "errors," where the "error" or residual is the...
4.0K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Friedman Two-way Analysis of Variance by Ranks
150
Friedman's Two-Way Analysis of Variance by Ranks is a nonparametric test designed to identify differences across multiple test attempts when traditional assumptions of normality and equal variances do not apply. Unlike conventional ANOVA, which requires normally distributed data with equal variances, Friedman's test is ideal for ordinal or non-normally distributed data, making it particularly useful for analyzing dependent samples, such as matched subjects over time or repeated measures...
150
Probability Histograms
11.1K
A probability histogram is a visual representation of a probability distribution. Similar a typical histogram, the probability histogram consists of contiguous (adjoining) boxes. It has both a horizontal axis and a vertical axis. The horizontal axis is labeled with what the data represents. The vertical axis is labeled with probability. Each rectangular bar in the histogram is 1 unit wide, which suggests that the area under each bar equals the probability, P(x), where x is 1, 2, 3, and so on.
11.1K
Skewness
10.9K
The measures of central tendency calculated from a data set may not reveal much about its intrinsic distribution. If a plot is made of the data set’s values, the mean and the median may not only differ, but also the plot may have more values on one side of the central tendencies. Such a data set is said to be skewed towards that side.
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
The longer the tail of the plot on one side, the more skewed it is. The skewness of a data set’s values suggests that the measures of central tendency...
10.9K


