scGCC:用于scRNA-Seq数据分析的图形对比集群与邻域增强
IEEE journal of biomedical and health informatics
|September 26, 2023
概括
我们介绍了scGCC,它是用于单细胞RNA测序 (scRNA-seq) 数据集群的图形自主监督对比学习模型. scGCC通过学习被拒绝的嵌入,提高聚类准确性和稳定性来增强细胞类型识别.
科学领域:
- 计算生物学 计算生物学
- 基因组学就是基因组学.
- 生物信息学是一种生物信息学.
背景情况:
- 单细胞RNA测序 (scRNA-seq) 对于分析细胞异质性至关重要.
- 在scRNA-seq数据中的细胞聚类对于识别细胞类型和亚型至关重要.
- scRNA-seq数据对集群提出了诸如高维度,稀疏性和批量效应等挑战.
研究的目的:
- 提出scGCC,一个新的图形自我监督对比式学习模型用于scRNA-seq数据集群.
- 解决scRNA-seq数据分析中的计算挑战,改善细胞类型识别.
- 提高scRNA-seq数据集中的细胞聚类的准确性和稳定性.
主要方法:
- 开发了scGCC,一个图形自主监督的对比学习模型,具有表示和聚类模块.
- 用于细胞表示学习和特征提取的图表注意力网络 (GAT).
- 实施了五种数据增强方法,以增加数据多样性并减少过度匹配.
主要成果:
- scGCC学习了有利于集群的低维的无效嵌入.
- 在14个现实世界scRNA-seq数据集中实现了非凡的准确性和稳定性.
- 通过下游任务 (如批量效应去除和轨迹推断) 证明了生物有效性.
结论:
- scGCC有效地解决了scRNA-seq数据集群中的挑战.
- 该模型改善了细胞类型的识别和新型亚型的发现.
- scGCC为scRNA-seq数据分析提供了一种强大而准确的方法.
更多相关视频
06:24Multiplexed Analysis of Retinal Gene Expression and Chromatin Accessibility Using scRNA-Seq and scATAC-Seq
Published on: March 12, 2021
3.6K
12:27Large-scale Reconstructions and Independent, Unbiased Clustering Based on Morphological Metrics to Classify Neurons in Selective Populations
Published on: February 15, 2017
7.0K
相关概念视频
RNA-seq
10.0K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
10.0K
Comparing Copy Number Variations and SNPs
17.7K
Sequencing of the human genome has opened up several best-kept secrets of the genome. Scientists have identified thousands of genome variations that exist within a population. These variations can be a single nucleotide or a larger chromosomal variation.
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
17.7K
Next-generation Sequencing
89.8K
The first human genome sequencing project cost $2.7 billion and was declared complete in 2003, after 15 years of international cooperation and collaboration between several research teams and funding agencies. Today, with the advent of next-generation sequencing technologies, the cost and time of sequencing a human genome have dropped over 100 fold.
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
Next-Generation Sequencing Methods
Although all next-generation methods use different technologies, they all share a set of standard features....
89.8K
Cluster Sampling Method
12.0K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
12.0K
Improving Translational Accuracy
11.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
11.5K
