探索数据缺失对机器学习公平性的不公平影响
Sitao Min1, Hafiz Asif2, Jaideep Vaidya1
1Rutgers University, Newark, NJ, 07102, USA.
概括
人工智能 (AI) 和机器学习 (ML) 模型中缺失的数据可以加剧公平差异,特别是当缺失与敏感属性相关时. 解决数据缺口对于公平的人工智能至关重要.
科学领域:
- 计算机科学 计算机科学
- 数据科学数据科学数据科学
- 人工智能的人工智能
背景情况:
- 数据驱动模型和AI/ML是社会决策的组成部分.
- 对算法公平性的担忧很大.
- 缺少数据对公平性的影响尽管普遍存在,但仍未得到充分研究.
研究的目的:
- 系统地评估缺少的数据如何影响分类器的公平性.
- 调查与受保护的类别和结果相关的缺失数据的作用.
主要方法:
- 分析了150个实验数据集变体,反映了现实世界的场景.
- 使用一个全面的框架,涵盖缺失的数据模式,速率和缓解策略.
主要成果:
- 缺少的数据,特别是与敏感的属性和结果相关联时,可能会加剧公平差异.
- 即使是少量的缺失也会对公平性产生重大影响.
- 系统性缺失对公平的人工智能构成重大挑战.
结论:
- 解决缺失的数据对于评估和确保算法公平性至关重要.
- 公平性评估必须考虑缺失数据的存在和模式.
- 缺失数据的缓解策略对于开发可靠的AI系统至关重要.
相关概念视频
Bias
5.1K
Bias refers to any tendency that prevents a question from being considered unprejudiced. In research, bias occurs when one outcome or answer is selected or encouraged over others in sampling or testing. Bias can occur during any research phase, including study design, data collection, analysis, and publication.
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
In statistics, a sampling bias is created when a sample is collected from a population, and some members of the population are not as likely to be chosen as others (remember, each member...
5.1K
Bias in Epidemiological Studies
700
Biases can arise at various stages of research, from study design and data collection to analysis and interpretation. Recognizing and addressing these biases is essential to ensure the validity and reliability of epidemiological findings.Broadly speaking, biases in epidemiology fall into three main categories: selection bias, information bias, and confounding. A more detailed description of possible biases is:
700
Censoring Survival Data
251
Survival analysis is a statistical method used to analyze time-to-event data, often employed in fields such as medicine, engineering, and social sciences. One of the key challenges in survival analysis is dealing with incomplete data, a phenomenon known as "censoring." Censoring occurs when the event of interest (such as death, relapse, or system failure) has not occurred for some individuals by the end of the study period or is otherwise unobservable, and it might have many different...
251
Detection of Gross Error: The Q Test
6.4K
When one or more data points appear far from the rest of the data, there is a need to determine whether they are outliers and whether they should be eliminated from the data set to ensure an accurate representation of the measured value. In many cases, outliers arise from gross errors (or human errors) and do not accurately reflect the underlying phenomenon. In some cases, however, these apparent outliers reflect true phenomenological differences. In these cases, we can use statistical methods...
6.4K
Weighted Mean
5.4K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
5.4K
Statistical Inference Techniques in Hypothesis Testing: Parametric Versus Nonparametric Data
217
Statistical inference techniques, paramount in hypothesis testing, differentiate into two broad categories: parametric and nonparametric statistics.
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
Parametric statistics, as the name suggests, assumes that data follow a specific distribution, often a normal distribution. This assumption enables robust hypothesis testing and estimation. Parametric methods, like the Student's t-test or Goodness-of-fit test, are frequently employed in biostatistics due to their robustness. For instance,...
217


