三方:通过更精确的分区来解决现实的噪音标签
Lida Yu1, Xuefeng Liang2,3, Chang Cao3
1School of Arts and Sciences, Beijing Normal University, Beijing 100875, China.
Sensors (Basel, Switzerland)
|September 19, 2025
概括
本研究介绍了Tripartite,这是一种新的方法来处理大规模数据集中错误标记的数据. 三方有效地识别和减轻噪音标签的影响,改善深度学习模型的性能.
科学领域:
- 机器学习 机器学习
- 计算机科学 计算机科学
- 人工智能的人工智能
背景情况:
- 深度学习模型很容易在大型数据集中过度匹配错误标记的数据.
- 现有的方法经常错误地将不确定的噪音样品归类为干净的,因为损失值低,降低了模型性能.
- 解决噪音标签对于提高深度学习模型可靠性至关重要.
研究的目的:
- 提出一种新的方法,三方,将训练数据分成不确定的,干净的和杂的子集.
- 通过准确识别不确定的噪音样本来提高清洁子集的质量.
- 通过利用干净样品和减轻噪音标签的影响来提高深度模型性能.
主要方法:
- 根据两个网络和给定的标签之间的预测不一致性,制定了三方数据分区策略.
- 分类数据分为三个子集:不确定,干净和杂.
- 应用低重量学习对不确定的样本和半监督学习对杂的样本.
主要成果:
- 三方显著提高了清洁数据子集的纯度.
- 拟议的方法有效地以更高的精度过出噪音样本.
- 实验结果显示,与基准和现实世界数据集的最先进方法相比,其性能优越.
结论:
- 三方提供了一种更精确的方法来处理深度学习中的噪音标签.
- 该方法通过更好地利用清洁数据和减少错误标记样本的负面影响来提高模型性能.
- 三方展示了强大的潜力,以提高深度学习模型在大规模,现实世界的数据集的稳定性.
相关概念视频
Extraction: Partition and Distribution Coefficients
4.6K
The distribution law or Nernst's distribution law is the law that governs the distribution of a solute between two immiscible solvents. This law, also known as the partition law, states that if a solute is added to the mixture of two immiscible solvents at a constant temperature, the solute is distributed between the two solvents in such a way that the ratio of solute concentrations in the solvents remains constant at equilibrium.
For extracting a solute from an aqueous phase into an...
For extracting a solute from an aqueous phase into an...
4.6K
Survival Tree
389
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
389
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K
Quantifying and Rejecting Outliers: The Grubbs Test
3.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.6K
¹H NMR: Interpreting Distorted and Overlapping Signals
1.5K
Spin systems where the difference in chemical shifts of the coupled nuclei is greater than ten times J are called first-order spin systems. These nuclei are weakly coupled, and their chemical shifts and coupling constant can generally be estimated from the well-separated signals in the spectrum.
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
As Δν decreases and the signals move closer, the doublets appear increasingly distorted. The intensities of the inner lines increase at the cost of those of the outer lines as the signals are...
1.5K


