在不平衡的基因表达分类中选择特征的边际加权强有力的分辨分数
Sheema Gul1, Dost Muhammad Khan1, Saeed Aldahmani2
1Department of Statistics, Abdul Wali Khan University, Mardan, Pakistan.
PloS one
|June 10, 2025
概括
一种新的特征选择方法 - - 边际加权稳健歧视性得分 (MW-RDS) - - 能够有效地处理高维不平衡基因表达数据. 通过放大少数阶级的影响力和减少冗余性,MW-RDS提高了分类准确性.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 高维基因表达数据给二进制分类带来了挑战.
- 现有的特征选择方法与类不平衡和特征冗余性作斗争.
研究的目的:
- 为高维不平衡数据提出一个强大的特征选择方法.
- 为了增强基因/特征的歧视力和阶级分离.
主要方法:
- 引入了边际加权强健歧视性得分 (MW-RDS).
- 整合了少数放大因子和特定类别的稳定性权重.
- 从支向量和L1规则化中使用了边缘权重.
主要成果:
- 在9个基因表达数据集上,MW-RDS在现有方法上表现出优越的性能.
- 使用随机森林,SVM和加权kNN分类器进行评估.
- 实现了提高准确性,灵敏度,特异性,F1得分和精度.
结论:
- 对于高维度不平衡问题,MW-RDS是一种强大而有效的特征选择方法.
- 该方法成功地解决了阶级不平衡和冗余问题.
- 优于基因表达数据分类中的传统方法.
相关概念视频
Quantifying and Rejecting Outliers: The Grubbs Test
1.5K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.5K
Weighted Mean
4.9K
While taking the arithmetic, geometric, or harmonic mean of a sample data set, equal importance is assigned to all the data points. However, all the values may not always be equally important in some data sets. An intrinsic bias might make it more important to give more weightage to specific values over others.
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
For example, consider the number of goals scored in the matches of a tournament. While computing the average number of goals scored in the tournament, it may be more important to...
4.9K
Expected Frequencies in Goodness-of-Fit Tests
2.5K
A goodness-of-fit test is conducted to determine whether the observed frequency values are statistically similar to the frequencies expected for the dataset. Suppose the expected frequencies for a dataset are equal such as when predicting the frequency of any number appearing when casting a die. In that case, the expected frequency is the ratio of the total number of observations (n) to the number of categories (k).
2.5K
Regression Toward the Mean
6.3K
Regression toward the mean (“RTM”) is a phenomenon in which extremely high or low values—for example, and individual’s blood pressure at a particular moment—appear closer to a group’s average upon remeasuring. Although this statistical peculiarity is the result of random error and chance, it has been problematic across various medical, scientific, financial and psychological applications. In particular, RTM, if not taken into account, can interfere when...
6.3K
Wilcoxon Signed-Ranks Test for Median of Single Population
104
The Wilcoxon signed-rank test for the median of a single population is a nonparametric test used to evaluate whether the median of a population differs from a specified value. Unlike parametric tests, it does not require data to follow a normal distribution, making it suitable for non-normal or small samples. The test begins by calculating the difference (d) between each observation and the hypothesized median. The absolute values of these differences are ranked in ascending order, with ties...
104
Sensitivity, Specificity, and Predicted Value
213
In healthcare diagnostics, laboratory tests play a crucial role in identifying and diagnosing a wide range of medical conditions. However, interpreting test results is not always straightforward. An abnormal test result does not always confirm the presence of a disease, just as a normal result does not guarantee its absence. To assess the reliability of these diagnostic tools, healthcare practitioners rely on two key statistical indicators: sensitivity and specificity.
Sensitivity is the...
Sensitivity is the...
213


