ScInfoVAE:使用变异自编码器和扩展的相互信息规范化的单细胞转录数据的可解释的维度缩小
Weiquan Pan1, Faning Long2, Jian Pan1
1School of Mathematics and Statistics, Yulin Normal University, Yulin, China.
BioData mining
|June 10, 2023
概括
ScInfoVAE是一种用于单细胞RNA测序 (scRNA-seq) 分析的新方法,可以有效地识别复杂组织中的细胞类型. 这种方法增强了数据解释,并提高了学习表征的质量,以获得更好的生物洞察力.
科学领域:
- 计算生物学 计算生物学
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
背景情况:
- 单细胞RNA测序 (scRNA-seq) 揭示了细胞异质性,并有助于细胞类型识别.
- 变异自编码器 (VAE) 显示了学习scRNA-seq特征表示的前景.
- 标准VAE可以忽略隐藏的变量,具有灵活的解码分布.
研究的目的:
- 介绍ScInfoVAE,一种用于scRNA-seq数据的新型尺寸缩小方法.
- 使用scRNA-seq数据改善复杂组织中的细胞类型识别.
- 提高scRNA-seq. 的低维表示的可解释性和质量.
主要方法:
- 基于相互信息变异自编码器 (InfoVAE) 开发了ScInfoVAE.
- 整合了联合InfoVAE深度模型与零膨胀负二项式模型.
- 重建了目标函数,以消除scRNA-seq数据,并学习高效的低维表示.
主要成果:
- 在15个真实scRNA-seq数据集中,ScInfoVAE展示了高集群性能.
- 使用模拟数据调查了特征提取解释性.
- 可视化证实ScinfoVAE的低维表示保留了本地和全球社区结构.
- 该模型显著改善了变化的后部质量.
结论:
- 在scRNA-seq数据中,ScInfoVAE是一种有效的尺寸缩小和细胞类型识别方法.
- 该方法提供了可靠和可解释的低维表示.
- ScInfoVAE推进了复杂组织scRNA-seq数据的分析.
相关概念视频
Cell Specific Gene Expression
13.6K
Multicellular organisms contain a variety of structurally and functionally distinct cell types, but the DNA in all the cells originated from the same parent cells. The differences in the cells can be attributed to the differential gene expression. Liver cells, whose functions include detoxification of blood, production of bile to metabolize fats, and synthesis of proteins essential for metabolism, must express a specific set of genes to perform their functions. Gene expression also varies with...
13.6K
Improving Translational Accuracy
11.7K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.7K
Variability: Analysis
162
Measures of variability are statistical metrics that reveal the dispersion pattern within a dataset. They are pivotal in biostatistics, providing insights into the heterogeneity within health and biological data. Variability signifies the degree to which data points diverge from one another, helping researchers understand the potential range of values and associated uncertainty within the data.
The range is a simple measure of variability, indicating the difference between the highest and...
The range is a simple measure of variability, indicating the difference between the highest and...
162
One-Compartment Open Model: Wagner-Nelson and Loo Riegelman Method for ka Estimation
587
This lesson introduces two critical methods in pharmacokinetics, the Wagner-Nelson and Loo-Riegelman methods, used for estimating the absorption rate constant (ka) for drugs administered via non-intravenous routes. The Wagner-Nelson method relates ka to the plasma concentration derived from the slope of a semilog percent unabsorbed time plot. However, it is limited to drugs with one-compartment kinetics and can be impacted by factors like gastrointestinal motility or enzymatic degradation.
On...
On...
587


