链混合特征选择算法基于改进的灰狼优化算法
Xiaotong Bai1, Yuefeng Zheng1, Yang Lu1
1School of Mathematics and Computer, Jilin Normal University, Siping, Jilin, China.
PloS one
|October 8, 2024
概括
一个新的混合特征选择算法,Tandem Maximum Kendall Minimum Chi-Square and ReliefF Improved Grey Wolf Optimization (TMKMCRIGWO),提高了分类准确度和尺寸缩小率. 这种算法可以提高分类准确度和尺寸缩小率. 这种方法在复杂的问题解决中优于单个算法.
科学领域:
- 机器学习 机器学习
- 数据挖掘 数据挖掘
- 生物信息学是一种生物信息学.
背景情况:
- 单一特征选择方法通常在有效性和性能方面存在局限性.
- 结合不同技术的混合方法可以克服这些限制.
- 需要先进的算法来改善复杂数据集中的特征选择.
研究的目的:
- 为了提出一种新的混合特征选择算法,TMKMCRIGWO.
- 为了提高分类准确性和尺寸缩小率.
- 证明拟议的算法在现有方法上的优越性.
主要方法:
- TMKMCRIGWO算法采用一个两阶段的过方法,使用最大肯德尔最小奇平方 (MKMC) 和ReliefF.
- 一个包装算法,一个改进的灰狼优化 (IGWO) 随机干扰因子,被用于最佳子集选择.
- 过器和包装方法的双重组合旨在实现强大的特征选择.
主要成果:
- 与其他算法相比,TMKMCRIGWO算法在20个数据集中实现了至少0.1%的平均分类准确度增加.
- 平均维度缩小率 (DRR) 达到24.76%,低维数据集的比例更高 (41.04%) 和高维数据集的比例更低 (0.33%).
- 该算法展示了改进的模型概括能力和性能.
结论:
- 与单一方法相比,提议的TMKMCRIGWO算法在特征选择方面提供了更高的性能.
- 混合战略有效地解决了复杂的问题,从而提高了分类准确性和显著的尺寸缩小.
- TMKMCRIGWO显示了提高机器学习模型性能和通用性的前景.
相关概念视频
Conservation of Small Populations
13.1K
Small population sizes put a species at extreme risk of extinction due to a lack of variation, and a consequent decrease in adaptability. This weakens the chances of survival under pressures such as climate change, competition from other species, or new diseases. Large populations are more likely to survive pressures such as these, as such populations are more likely to harbor individuals that have genetic variants that are adaptive under new stresses. Small populations are much less...
13.1K
Hybrid Zones
16.9K
Hybrid zones are narrow regions where two closely related species interact, mate, and produce hybrids. Relative to either parent species, hybrids may possess distinct phenotypic or genetic differences that impact their survival and reproductive success. The genetic variances introduced by hybridization influence species diversity and speciation processes within the hybrid zone.
16.9K
Wald-Wolfowitz Runs Test I
614
The Wald-Wolfowitz test, also known as the runs test, is a nonparametric statistical test used to assess the randomness of a sequence of two different types of elements (e.g., positive/negative values, successes/failures). It examines whether the order of the elements in a sequence is random or if there is a pattern or trend present. This nonparametric test applies to any ordered data despite the population and sample data distribution, even if a higher sample size is available.
The test works...
The test works...
614
Wald-Wolfowitz Runs Test II
190
The Wald-Wolfowitz runs test, commonly referred to as the runs test, is a nonparametric test used to assess the randomness of ordered data. The test evaluates the number of runs, which are consecutive sequences of similar elements within the data. If the number of runs is significantly higher or lower than expected, the data is considered non-random, indicating a detectable pattern or structure.
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
For binary data, runs are identified using symbols such as + and −, or equivalently, 1s and...
190
Heuristics
78
Heuristics are problem-solving strategies that use mental shortcuts to simplify decision-making. Unlike algorithms, which must be followed precisely to achieve a correct result, heuristics offer a general problem-solving framework. They save time and energy but can sometimes lead to less rational decisions.
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
People often rely on heuristics when faced with an overload of information, limited time, low importance of the decision, limited information, or when a heuristic readily comes to mind. For...
78
Improving Translational Accuracy
9.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.3K


