带有 Universum 数据的颗粒球双支向量机
1Department of Computer Science and Engineering, Indian Institute of Technology Ropar, Rupnagar, 140001, Punjab, India.
概括
带有宇宙数据的颗粒球双支持向量机 (GBU-TSVM) 提高了分类的准确性和稳定性. 这种新的方法将数据模拟为超球, 提高噪音数据集的性能, 并优于现有方法.
科学领域:
- 机器学习
- 数据挖掘
- 模式识别
背景情况:
- 支持矢量机器 (SVM) 经常与有限的标记数据作斗争,并且对噪声和异常值敏感.
- 传统的双支持向量机 (TSVM) 将数据表示为点,从而限制了它们的稳定性和效率.
- 现有的方法缺乏有效的策略来处理噪音数据和利用未标记或类外信息.
研究的目的:
- 作为一个强大的分类框架,引入带有宇宙数据的颗粒球双支持向量机 (GBU-TSVM).
- 通过集成颗粒球计算和Universum数据来提高TSVM的性能.
- 提高分类准确性和计算效率,特别是在噪音和有限的标记数据的情况下.
主要方法:
- 将数据实例建模为超球,而不是TSVM框架中的点.
- 使用颗粒球计算来有效地分组数据并降低处理复杂度.
- 整合Universum数据 (目标类之外的样本) 来完善决策边界并提高概括性.
主要成果:
- 在最佳条件下,GBU-TSVM对Molec Biol Promoter数据集的准确度达到了92. 38%.
- 即使在20%的噪音污染下,该模型也保持了89.17%的准确性,显示出显著的稳定性.
- 在实验中,GBU-TSVM的表现始终优于基线模型,包括GBSVM,TSVM,GBTSVM,Pin-GTSVM和UTSVM.
结论:
- GBU-TSVM为具有挑战性的数据环境提供了卓越而强大的分类框架.
- 集成颗粒球计算和Universum数据显著提高了SVM的性能.
- 这种方法为开发更具弹性和准确的机器学习模型提供了有希望的方向.
相关概念视频
Classification of Systems-II
240
Continuous-time systems have continuous input and output signals, with time measured continuously. These systems are generally defined by differential or algebraic equations. For instance, in an RC circuit, the relationship between input and output voltage is expressed through a differential equation derived from Ohm's law and the capacitor relation,
240
Aggregates Classification
380
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
380
End Point Prediction: Gran Plot
580
A Gran plot is used to predict the equivalence volume or endpoint of a potentiometric or acid-base titration without reaching the endpoint. Typically, titration data is collected as a function of the titrant's volume up to a point less than the equivalence volume and then transformed into a linear format. The straight line is extended to the x-axis, indicating the necessary titrant volume to achieve the equivalence point.
For potentiometric titration, the Gran plot is created by plotting...
For potentiometric titration, the Gran plot is created by plotting...
580
Sequence Networks of Rotating Machines
140
A Y-connected synchronous generator, grounded through a neutral impedance, is designed to produce balanced internal phase voltages with only positive-sequence components. The generator's sequence networks include a source voltage that is exclusively in the positive-sequence network. The sequence components of line-to-ground voltages at the generator terminals illustrate this configuration.
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
Zero-sequence current induces a voltage drop across the generator's neutral impedance and other...
140
Classification of Systems-I
294
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
294
Quantifying and Rejecting Outliers: The Grubbs Test
2.0K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
2.0K


