一个基于集群的SMOTE双采样 (CSBBoost) 整体算法用于分类不平衡数据.
Amir Reza Salehi1, Majid Khedmati2
1Department of Industrial Engineering, Sharif University of Technology, 9414 Azadi Ave, P.O. Box 11155, Tehran, 1458889694, Iran.
Scientific reports
|March 2, 2024
概括
一个新的基于集群的合成少数群体过量采样技术 (SMOTE) 两种采样 (CSBBoost) 组合算法有效地分类不平衡的数据. 这种新的方法在平衡数据集和提高分类准确性方面优于现有方法.
科学领域:
- 机器学习 机器学习
- 数据科学数据科学数据科学
- 计算机科学 计算机科学
背景情况:
- 不平衡的数据集在机器学习分类任务中带来了重大挑战.
- 传统的方法经常遭受信息丢失或数据冗余的困扰.
- 在许多现实应用中,对少数群体阶级的准确分类至关重要.
研究的目的:
- 提出一种新的整体算法,以集群为基础的合成少数人过量采样技术 (SMOTE) 两者采样 (CSBBoost),以改进不平衡数据分类.
- 解决现有的过量采样和不足采样技术的局限性.
- 在不平衡的数据集上增强整体方法的性能.
主要方法:
- 拟议的CSBBoost算法集成了过量采样,不足采样和组合技术 (极端梯度提升,随机森林,袋装).
- 它旨在减轻因过量采样而导致的冗余和因样本不足而导致的信息丢失.
- 该算法采用基于集群的方法来生成合成数据.
主要成果:
- 与最先进的竞争算法相比,CSBBoost算法表现出明显优异的性能.
- 对20个基准不平衡数据集的性能进行了评估,使用F1分数和接收器操作特征曲线 (AUC) 下的区域.
- 提出的方法在不平衡的数据上实现了更好的分类准确性和稳定性.
结论:
- CSBBoost算法为分类不平衡数据提供了一个有效的解决方案.
- 它成功地平衡了数据集,同时保留了关键信息.
- 算法的适用性通过真实世界的数据集分析来验证.
相关概念视频
Classification of Systems-I
186
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
186
Classification of Systems-II
146
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
146
Aggregates Classification
321
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
321
Classification of Signals
460
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
460
Prediction Intervals
2.3K
The interval estimate of any variable is known as the prediction interval. It helps decide if a point estimate is dependable.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
2.3K
Classification of Leukocytes
1.9K
Leukocytes are classified into two groups based on the presence or absence of cytoplasmic granules. Granular leukocytes, which contain granules, belong to the myeloid lineage and are divided into three subtypes: neutrophils, eosinophils, and basophils. These cells are roughly spherical and characterized by the granules in their cytoplasm.
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
1.9K


