伪标签引导的结构歧视子空间学习用于无监督的特征选择
IEEE transactions on neural networks and learning systems
|October 5, 2023
概括
本研究介绍了伪标签引导的结构歧视子空间学习 (PSDSL),这是一种新的无监督特征选择方法. 关注DSL统一了概率图形构造和伪标签学习,以提高数据集群性能.
科学领域:
- 机器学习 机器学习
- 数据科学数据科学数据科学
- 生物信息学是一种生物信息学.
背景情况:
- 传统的特征选择方法经常独立执行阶段,限制了适应性.
- 现有的方法面临着稀疏性和参数调整的挑战,特别是在使用L1-norm.
- 无监督的特征选择对于有效的数据聚类和分析至关重要.
研究的目的:
- 为了提出一种新的无监督特征选择方法,P উদ্বেগDSL.
- 将概率图结构和特征选择统一到一个框架中.
- 加强特征歧视和稳定性,以改善下游集群任务.
主要方法:
- 在适应性学习的特征选择中引入了概率图形构造.
- 开发了一个伪标签引导学习机制.
- 结合基于图形的方法,最大限度地利用痕迹比来实现类间分散.
- 采用L0-规范约束对行稀疏性和特征稳定性,解决L1-规范限制.
主要成果:
- 在九个现实世界数据集上证明了P উদ্বেগDSL的有效性.
- 在三个生物单细胞RNA测序 (ScRNA-seq) 基因数据集上验证了该方法的性能.
- 与现有方法相比,取得了更好的数据聚类结果.
结论:
- 关注DSL提供了一个统一的和适应性框架,用于无监督的功能选择.
- 该方法有效地改善了特征歧视和稳定性.
- 关注DSL显示了在各种应用中增强数据聚类的重大前景,包括生物信息学.
相关概念视频
Structural Classification of Joints
3.5K
Joints, also known as articulations, are classified based on their structural characteristics, i.e., based on whether the articulating surfaces of the adjacent bones are directly connected by fibrous connective tissue or cartilage, or whether the articulating surfaces contact each other within a fluid-filled joint cavity. These differences serve to divide the joints of the body into three structural classifications.
A fibrous joint is where the adjacent bones are united by fibrous connective...
A fibrous joint is where the adjacent bones are united by fibrous connective...
3.5K
Quantifying and Rejecting Outliers: The Grubbs Test
1.6K
Sometimes, a data set can have a recorded numerical observation that greatly deviates from the rest of the data. Assuming that the data is normally distributed, a statistical method called the Grubbs test can be used to determine whether the observation is truly an outlier. To perform a two-tailed Grubbs test, first, calculate the absolute difference between the outlier and the mean. Then, calculate the ratio between this difference and the standard deviation of the sample. This...
1.6K
Residuals and Least-Squares Property
7.4K
The vertical distance between the actual value of y and the estimated value of y. In other words, it measures the vertical distance between the actual data point and the predicted point on the line
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
If the observed data point lies above the line, the residual is positive, and the line underestimates the actual data value for y. If the observed data point lies below the line, the residual is negative, and the line overestimates the actual data value for y.
The process of fitting the best-fit...
7.4K


