在变异自编码器中利用相互信息来改善单细胞RNA测序数据的维度减少:scInfoMaxVAE方法
Pham Nhat Duy1, Nguyen Phuong Thao1, Thanh Le1
1Faculty of Information Technology, University of Science, Ho Chi Minh City, 700000, Viet Nam; Viet Nam National University, Ho Chi Minh City, 720325, Viet Nam.
Computational biology and chemistry
|August 29, 2025
概括
scInfoMaxVAE是一种用于单细胞RNA测序分析的新工具,通过最大限度地提高相互信息和处理技术噪音来改善数据表示. 它为复杂的生物数据提供了强大的维度缩小和细胞类型分类.
科学领域:
- 计算生物学
- 基因组学
- 生物信息学
背景情况:
- 单细胞RNA测序 (scRNA-seq) 可以产生高维度,稀疏的数据.
- 技术噪音和稀疏性使得数据的准确表示和解释成为一个挑战.
- 现有的方法通常在多种scRNA-seq数据集中难以稳定.
研究的目的:
- 为scRNA-seq数据开发一种新的变异自编码模型.
- 增强维度缩小和细胞类型分类能力.
- 解决scRNA-seq数据表示中的稀疏性和技术噪音.
主要方法:
- 推出了scInfoMaxVAE,一个相互信息最大化的变量自动编码器.
- 纳入为scRNA-seq数据量身定制的零膨胀计数概率.
- 通过使用统一的质量控制和注释管道对12个公共scRNA-seq数据集进行评估.
主要成果:
- scInfoMaxVAE在不同的数据集中展示了竞争性的集群和结构保存.
- 达到了高标准化的相互信息 (NMI) 评分 (0.94),符合最先进的方法.
- 与scVI和t-SNE相比,同质性 (0. 89) 和调整后的兰德指数显著改善.
结论:
- scInfoMaxVAE为scRNA-seq表示学习提供了强大且可重复的方法.
- 它的信息理论培训和零通货膨胀模型提高了对异质数据的性能.
- 在scRNA-seq工作流程中提供了一个有希望的替代方案来减少维度和细胞类型分类.
相关概念视频
RNA-seq
10.4K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.4K
Improving Translational Accuracy
11.8K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.8K
Variance
10.5K
The deviations show how spread out the data are about the mean. A positive deviation occurs when the data value exceeds the mean, whereas a negative deviation occurs when the data value is less than the mean. If the deviations are added, the sum is always zero. So one cannot simply add the deviations to get the data spread. By squaring the deviations, the numbers are made positive; thus, their sum will also be positive.
The standard deviation measures the spread in the same units as the...
The standard deviation measures the spread in the same units as the...
10.5K


