垃圾邮件分类使用支持向量机器使用范德瓦尔登等级分数注意力分数
1School of Mathematics and Statistics, Xiamen University of Technology.
Journal of visualized experiments : JoVE
|November 17, 2025
概括
一个新的Van der Waerden排名得分功能增强了注意力支持向量机 (VWR-Attn-SVM) 提供了高效的垃圾邮件分类. 这种方法提高了准确性,并降低了计算成本,以提高网络安全性.
科学领域:
- 计算机科学 计算机科学
- 机器学习 机器学习
- 网络安全 网络安全
背景情况:
- 垃圾邮件对网络安全和通信效率构成重大威胁.
- 传统的垃圾邮件检测方法,如传统的机器学习和深度学习,在处理高维数据或计算资源方面存在局限性.
研究的目的:
- 引入一个高效和可解释的垃圾邮件分类方法.
- 解决现有的垃圾邮件检测技术的局限性.
主要方法:
- 开发了一个新的范德瓦尔登排名得分功能,增强了注意力支持向量机 (VWR-Attn-SVM).
- 范德瓦尔登等级转换用于文本特征正常化,增强异常值的稳定性并保留顺序关系.
- 为了优化特征选择,采用了具有非线性处理和规范化的增强注意力机制.
主要成果:
- 在UCI Spambase和印度尼西亚垃圾邮件数据集上,VWR-Attn-SVM在传统分类器上表现优越.
- 该方法实现了更高的准确性,精度,回忆,F1得分和AUC.
- 该方法将高性能与降低的计算成本相结合.
结论:
- VWR-Attn-SVM为垃圾邮件分类提供了一个高效和可解释的解决方案.
- 该方法显示了在其他基于文本的平台 (如消息和社交媒体) 中的应用潜力.
- 这种技术在有效打击垃圾邮件方面提供了有希望的进步.
相关概念视频
Classification of Signals
1.3K
In signal processing, signals are classified based on various characteristics: continuous-time versus discrete-time, periodic versus aperiodic, analog versus digital, and causal versus noncausal. Each category highlights distinct properties crucial for understanding and manipulating signals.
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
1.3K
Force Classification
2.3K
Forces play a crucial role in the study of physics and engineering. They are essential in describing the motion, behavior, and equilibrium of objects in the physical world. Forces can be classified based on their origin, type, and direction of action.
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
Contact and non-contact forces are two of the most widely used categories of forces. As the name suggests, contact forces require physical contact between two objects to act upon each other. Examples of contact forces include frictional,...
2.3K
Classification of Systems-I
543
Linearity is a system property characterized by a direct input-output relationship, combining homogeneity and additivity.
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
543
Quantifying and Rejecting Outliers: The Grubbs Test
3.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
3.5K
Classification of Leukocytes
4.9K
Leukocytes are classified into two groups based on the presence or absence of cytoplasmic granules. Granular leukocytes, which contain granules, belong to the myeloid lineage and are divided into three subtypes: neutrophils, eosinophils, and basophils. These cells are roughly spherical and characterized by the granules in their cytoplasm.
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
Neutrophils are the most abundant type of granular leukocytes, comprising 50-70% of all leukocytes. They feature small, evenly distributed granules and a...
4.9K
Ranks
448
Unlike parametric methods, nonparametric statistics are ideal for nominal and ordinal data, requiring fewer assumptions about the population's nature or distribution. This makes nonparametric methods easier to apply and interpret, as they do not depend on parameters like mean or standard deviation. One common approach in nonparametric analysis is to sort data according to a specific criterion. For instance, we might arrange weather data from hottest to coldest days in a month or rank cities...
448


