Related Experiment Video
Updated: Jan 19, 2026

Single-cell RNA Sequencing and Analysis of Human Pancreatic Islets
Published on: July 18, 2019
Evaluation of methods to assign cell type labels to cell clusters from single-cell RNA-sequencing data
J Javier Diaz-Mejia1,2,3, Elaine C Meng3, Alexander R Pico4
1Princess Margaret Cancer Centre, University Health Network, Toronto, ON, M5G 2M9, Canada.
Abstract:
Background: Identification of cell type subpopulations from complex cell mixtures using single-cell RNA-sequencing (scRNA-seq) data includes automated steps from normalization to cell clustering. However, assigning cell type labels to cell clusters is often conducted manually, resulting in limited documentation, low reproducibility and uncontrolled vocabularies. This is partially due to the scarcity of reference cell type signatures and because some methods support limited cell type signatures. Methods: In this study, we benchmarked five methods representing first-generation enrichment analysis (ORA), second-generation approaches (GSEA and GSVA), machine learning tools (CIBERSORT) and network-based neighbor voting (METANEIGHBOR), for the task of assigning cell type labels to cell clusters from scRNA-seq data. We used five scRNA-seq datasets: human liver, 11 Tabula Muris mouse tissues, two human peripheral blood mononuclear cell datasets, and mouse retinal neurons, for which reference cell type signatures were available. The datasets span Drop-seq, 10X Chromium and Seq-Well technologies and range in size from ~3,700 to ~68,000 cells. Results: Our results show that, in general, all five methods perform well in the task as evaluated by receiver operating characteristic curve analysis (average area under the curve (AUC) = 0.91, sd = 0.06), whereas precision-recall analyses show a wide variation depending on the method and dataset (average AUC = 0.53, sd = 0.24). We observed an influence of the number of genes in cell type signatures on performance, with smaller signatures leading more frequently to incorrect results. Conclusions: GSVA was the overall top performer and was more robust in cell type signature subsampling simulations, although different methods performed well using different datasets. METANEIGHBOR and GSVA were the fastest methods. CIBERSORT and METANEIGHBOR were more influenced than the other methods by analyses including only expected cell types. We provide an extensible framework that can be used to evaluate other methods and datasets at https://github.com/jdime/scRNAseq_cell_cluster_labeling.
More Related Videos
Related Concept Videos
11:34Single-cell RNA Sequencing and Analysis of Human Pancreatic Islets
06:36Gut Isolation from Zebrafish Larvae for Single-cell RNA Sequencing
08:53Isolation and Transcriptome Analysis of Plant Cell Types
11:26Single-cell RNA-Seq of Defined Subsets of Retinal Ganglion Cells
05:54Preparation of a Non-Cardiomyocyte Cell Suspension for Single-Cell RNA Sequencing from a Post-Myocardial Infarction Adult Mouse Heart
07:49Single-cell RNA Sequencing of Fluorescently Labeled Mouse Neurons Using Manual Sorting and Double In Vitro Transcription with Absolute Counts Sequencing (DIVA-Seq)

