智能适应组合模型用于多类不平衡的非静止数据流
Abdul Sattar Palli1,2, Jafreezal Jaafar3,4, Mohamad Hanif Md Saad5
1Department of Computer and Information Sciences, Universiti Teknologi PETRONAS, 32610, Seri Iskandar, Perak Darul Ridzuan, Malaysia. abdulsattarpalli@gmail.com.
Scientific reports
|July 2, 2025
概括
本研究介绍了智能自适应组合模型 (SAEM),以解决多类数据流中的概念漂移和类不平衡. 通过适应数据变化和重新加权少数阶级,SAEM显著提高了在线机器学习模型的性能.
科学领域:
- 机器学习 机器学习
- 数据科学数据科学数据科学
- 人工智能的人工智能
背景情况:
- 实时流数据通常同时呈现概念漂移和类不平衡,降低在线机器学习模型的性能.
- 现有的解决方案主要在二进制类流中解决这些问题,仅限于多类场景.
- 集体学习,一种共同的方法,当新的分类器没有接受反映概念变化的相关数据的培训时,就会遇到困难.
研究的目的:
- 提出一种新的智能自适应合组模型 (SAEM),以有效处理多类数据流中的概念漂移和类不平衡.
- 在动态环境中增强在线机器学习模型的稳定性和准确性.
- 解决当前集合方法在适应不断变化的数据概念方面的局限性.
主要方法:
- SAEM监控特征级数据分布变化,以识别概念漂移.
- 一个背景组合被用来训练新的分类器对显示变化的数据进行分类.
- 动态类失衡比重应用于少数类实例以减轻失衡.
主要成果:
- 拟议的SAEM在八个不同的数据流中,与最先进的方法相比,表现优越.
- 观察到显著的平均改善:15.86%的准确性,20.35%的卡帕,16.12%的F1得分,15.58%的精度和16.42%的回忆.
- 使用弗里德曼测试的统计分析证实了关键指标的显著性能差异.
结论:
- SAEM为在线学习应用程序提供了有效和高效的解决方案,这些应用程序在多类数据中面临概念漂移和类不平衡.
- 该模型的自适应性和处理不平衡数据有助于提高其性能.
- 这些发现支持SAEM在动态实时数据环境中保持高模型性能的能力.
相关概念视频
Classification of Systems-I
319
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
319
Classification of Signals
915
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
915
Aggregates Classification
389
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
389
Survival Tree
166
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
Building a Survival Tree
Constructing a...
Building a Survival Tree
Constructing a...
166
Classification of Systems-II
242
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
242
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
131
Pharmacokinetic models are mathematical constructs that represent and predict the time course of drug concentrations in the body, providing meaningful pharmacokinetic parameters. These models are categorized into compartment, physiological, and distributed parameter models.
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
131


