学生损失:在不准确的监督中向概率假设迈进
IEEE transactions on pattern analysis and machine intelligence
|January 23, 2024
概括
这项研究介绍了学生损失,这是机器学习中处理噪音标签的新方法. 通过对学生分布的深度特征进行建模,它可以有效地区分干净的数据和错误标记的数据,从而提高学习的稳定性.
科学领域:
- 机器学习 机器学习
- 计算机科学 计算机科学
- 人工智能的人工智能
背景情况:
- 使用噪音标签的学习在数据集策划中构成了重大挑战.
- 现有的方法往往不分青红白地处理错误标记和干净的样品,从而限制了强度.
- 清洁和错误标记的数据之间的自然差异往往被忽视.
研究的目的:
- 开发一种新的方法,在有噪音标签的情况下提高学习稳定性.
- 为了利用学生分布的特性来进行数据选择和抗噪声.
- 在不准确的监督场景中引入一个指标学习策略,以提高在不准确的监督场景中的表现.
主要方法:
- 提出了一个新的损失函数,称为学生损失,基于假设相同标签的深度特征遵循学生分布.
- 将学生分布嵌入到学习过程中,以利用其曲线的度来选择数据.
- 通过结合一套度量学习策略,开发出一个大幅度的学生 (LT) 损失.
- 引入了一种新的方法,在特征表示中使用先前的概率假设,以减少错误标记样本的影响.
主要成果:
- 学生损失方法证明了自然的数据选择性,导致干净的样本紧密聚合,错误标记的样本分散.
- 拟议的LT损失显著提高了抵抗错误标签样本的能力.
- 这种方法有效地减少了错误标记样本的贡献,甚至超过了现有的强损失.
- 实验显示了性能的大幅提高,在某些条件下超过50%,特别是在不准确的监督下.
结论:
- 学生损失框架通过建模特征分布,为学习有噪音标签提供了有效的策略.
- LT损失为杂的标签学习提供了强大的增强,超过了最先进的方法.
- 这项工作开创了先验概率假设在特征表示中用于机器学习中的降噪的使用.
更多相关视频
06:45Task Interruption and Resumption Paradigm for Testing the Activation and Pursuit of an Abstract Thinking Goal
Published on: April 18, 2017
6.2K
10:26Problem-Solving Before Instruction PS-I: A Protocol for Assessment and Intervention in Students with Different Abilities
Published on: September 11, 2021
4.0K
相关概念视频
Hindsight Biases
3.4K
Hindsight bias leads you to believe that the event you just experienced was predictable, even though it really wasn’t. In other words, you knew all along that things would turn out the way they did. Can you relate this to the phrase "Hindsight is 20/20" now?
3.4K
Estimating Population Mean with Unknown Standard Deviation
7.7K
In practice, we rarely know the population standard deviation. In the past, when the sample size was large, this did not present a problem to statisticians. They used the sample standard deviation s as an estimate for σ and proceeded as before to calculate a confidence interval with close enough results. However, statisticians ran into problems when the sample size was small. A small sample size caused inaccuracies in the confidence interval.
William S. Gosset (1876–1937) of the...
William S. Gosset (1876–1937) of the...
7.7K
Accuracy and Errors in Hypothesis Testing
200
Hypothesis testing is a fundamental statistical tool that begins with the assumption that the null hypothesis H0 is true. During this process, two types of errors can occur: Type I and Type II. A Type I error refers to the incorrect rejection of a true null hypothesis, while a Type II error involves the failure to reject a false null hypothesis.
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
In hypothesis testing, the probability of making a Type I error, denoted as α, is commonly set at 0.05. This significance level indicates a 5%...
200
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Assumptions of Survival Analysis
128
Survival models analyze the time until one or more events occur, such as death in biological organisms or failure in mechanical systems. These models are widely used across fields like medicine, biology, engineering, and public health to study time-to-event phenomena. To ensure accurate results, survival analysis relies on key assumptions and careful study design.
128
Reliability and Validity
12.7K
Reliability and validity are two important considerations that must be made with any type of data collection. Reliability refers to the ability to consistently produce a given result. In the context of psychological research, this would mean that any instruments or tools used to collect data do so in consistent, reproducible ways.
12.7K
