通过边界增强策略的元学习不平衡分类框架与贝叶斯不平衡影响指数
Qiangwei Li1, Xin Gao1, Heping Lu2
1School of Artificial Intelligence, Beijing University of Posts and Telecommunications, Beijing, 100876, China.
概括
这项研究引入了一种用于不平衡分类的新型元学习框架,增强边界样本,以提高分类器在不同数据集中的适应性和性能,而无需进行参数调整.
科学领域:
- 机器学习 机器学习
- 数据科学数据科学数据科学
- 人工智能的人工智能
背景情况:
- 不平衡的分类由于不同的数据集特征带来了挑战,阻碍了算法级方法的最佳参数选择.
- 由于数据集特定的特征,如不平衡比率和数据维度,现有的方法往往难以实现普遍性.
研究的目的:
- 为不平衡的分类提出一个普遍的元学习框架,克服参数调整的困难.
- 在微调阶段增强分类器的适应性和决策稳定性.
主要方法:
- 引入了一个具有多个培训任务的元学习框架,每个任务都包括用于元分类器培训和优化的支持和查询集.
- 开发了一个边界增强策略,利用贝叶斯失衡影响指数 (IBI3) 来识别和插曲边界样本.
- 在边界样本上应用特征插值,以加强分类边界并减轻决策偏差.
主要成果:
- 与24种传统的不平衡分类技术相比,提出的方法显示出更高的性能.
- 在38个不同的不平衡的公共数据集中获得了高的F-测量和G-平均分数.
- 具有强大的通用性,在不需要参数调整的情况下有效执行.
结论:
- 具有边界增强的元学习框架为不平衡的分类提供了强大的和通用的解决方案.
- 贝叶斯失衡影响指数 (IBI3) 有效地识别了针对目标增强的边界样本.
- 这种方法在不平衡的数据集中显著减轻了决策偏见,从而改善了概括性.
相关概念视频
Weighted Mean
4.9K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
4.9K
Prediction Intervals
2.2K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.2K
Classification of Systems-I
168
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
168
Classification of Systems-II
133
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
133
Aggregates Classification
300
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
300
Strategies for Assessing and Addressing Confounding
82
Confounding is a critical issue in epidemiological studies, often leading to misleading conclusions about associations between exposures and outcomes. It occurs when the relationship between the exposure and the outcome is mixed with the effects of other factors that influence the outcome. Given that, addressing confounding is of high importance for drawing accurate inferences in research.
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
Confounding can be addressed at both the design phase of a study and through analytical methods after data...
82
