scSFCL:Deep clustering of scRNA-seq data with subspace feature confidence learning
Xiaokun Meng1, Yuanyuan Zhang1, Xiaoyu Xu1
1School of Information and Control Engineering, Qingdao University of Technology, Qingdao, Shandong 266520, China.
Computational Biology and Chemistry
|November 26, 2024
Summary
A new deep clustering method, scSFCL, enhances single-cell RNA sequencing (scRNA-seq) data analysis. It improves cell type diversity identification by learning feature confidence and fusing structural information for more accurate clustering.
Area of Science:
- Bioinformatics
- Computational Biology
- Genomics
Background:
- Single-cell RNA sequencing (scRNA-seq) enables gene expression analysis for cell type diversity.
- High dimensionality, sparsity, and noise in scRNA-seq data challenge traditional clustering methods.
- Existing methods struggle to fully utilize discriminative attribute information and accurately capture cell diversity.
Purpose of the Study:
- To propose a novel deep clustering method, scSFCL, for scRNA-seq data.
- To address the limitations of traditional clustering in handling complex single-cell data.
- To improve the accuracy and effectiveness of cell type identification from scRNA-seq data.
Main Methods:
- Developed scSFCL based on subspace feature confidence learning.
- Utilized kernel density to divide subspaces and filter discriminative feature subsets.
- Employed a graph convolutional network (GCN) with weighting for feature confidence learning.
- Integrated GCN and a denoising variational autoencoder based on zero-inflated negative binomials (DVAE-ZINB) for mutually supervised clustering.
Main Results:
- scSFCL demonstrated significantly improved clustering performance on multiple scRNA-seq datasets.
- The method effectively filters discriminative feature subsets and learns their confidence.
- Complementary fusion of structural and idiosyncratic information enhanced clustering accuracy.
Conclusions:
- scSFCL provides an effective solution for deep clustering of scRNA-seq data.
- The proposed method overcomes limitations of traditional approaches in capturing cell type diversity.
- Enhanced feature learning and information fusion contribute to superior clustering performance.
More Related Videos
Related Concept Videos
Cluster Sampling Method
11.6K
Appropriate sampling methods ensure that samples are drawn without bias and accurately represent the population. Because measuring the entire population in a study is not practical, researchers use samples to represent the population of interest.
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
To choose a cluster sample, divide the population into clusters (groups) and then randomly select some of the clusters. All the members from these clusters are in the cluster sample. For example, if you randomly sample four departments from your...
11.6K
RNA-seq
9.8K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.8K
Confidence Coefficient
7.5K
The confidence coefficient is also known as the confidence level or degree of confidence. It is the percent expression for the probability, 1-α, that the confidence interval contains the true population parameter assuming that the confidence interval is obtained after sufficient unbiased sampling; for example, if the CL = 90%, then in 90 out of 100 samples the interval estimate will enclose the true population parameter. Here α is the area under the curve, distributed equally under...
7.5K
Confocal Fluorescence Microscopy
13.1K
Confocal microscopy is an advanced microscopic technique. The prime advantage of the confocal microscope over other microscopy techniques is its ability to block the out-of-focus light from the illuminated samples using pinholes. It is widely used with fluorescence optics to obtain high-resolution, sharp contrast images. Unlike optical microscopes, confocal microscopes use a focused beam of light laser to scan the entire sample surface at different z-planes. These microscopes are, therefore,...
13.1K
Improving Translational Accuracy
9.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
9.1K
Aggregates Classification
305
Aggregate classification is generally based on its size, petrographic characteristics, weight, and source. Size classification ranges from coarse to fine aggregates, defined by the size of the particles. Coarse aggregates are particles that do not pass through ASTM sieve No. 4, and aggregates that pass through the sieve are fine aggregates.
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
305


