Related Experiment Video
Updated: Sep 8, 2025

10:12
Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
Published on: January 10, 2019
18.7K
Evaluating the performance of dropout imputation and clustering methods for single-cell RNA sequencing data
Junlin Xu1, Lingyu Cui2, Jujuan Zhuang3
1College of Computer Science and Electronic Engineering, Hunan University, Changsha, Hunan, 410082, China.
Computers in Biology and Medicine
|June 13, 2022
Summary
Choosing the right imputation and clustering methods is crucial for accurate single-cell RNA sequencing (scRNA-seq) analysis. Performance varies by dataset size, with different combinations excelling on small versus large datasets.
Area of Science:
- Computational Biology
- Bioinformatics
- Genomics
Background:
- Single-cell RNA sequencing (scRNA-seq) enables transcriptome analysis at single-cell resolution.
- Cell clustering is vital for identifying cell subtypes and inferring lineage from scRNA-seq data.
- Technical noise leads to zero counts (dropout), complicating accurate clustering.
Purpose of the Study:
- To evaluate the impact of dropout imputation methods on scRNA-seq clustering performance.
- To identify optimal combinations of imputation and clustering algorithms for scRNA-seq data.
- To assess how dataset size influences the performance of imputation-clustering strategies.
Main Methods:
- Assessed nine dropout imputation methods combined with eight clustering algorithms.
- Utilized 10 well-annotated scRNA-seq datasets of varying sample sizes.
- Evaluated clustering performance and data visualization quality (t-SNE).
Main Results:
- Imputation methods generally enhance clustering performance and data visualization quality.
- Optimal imputation-clustering combinations are dataset-size dependent.
- Single-cell analysis via expression recovery (imputation) + Sparse Subspace Clustering (SSC) performs well on smaller datasets.
- Adaptively-thresholded low-rank approximation (imputation) + single-cell interpretation via multikernel learning (SIMLR) excels on larger datasets.
Conclusions:
- Dropout imputation is beneficial for improving scRNA-seq clustering accuracy.
- Algorithm selection for scRNA-seq analysis should consider dataset size.
- Specific imputation-clustering pairs offer superior performance depending on data scale.

