Related Experiment Video
Updated: Jan 9, 2026

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
Published on: January 10, 2019
CANTAO: guiding clustering and annotation in single-cell RNA sequencing using average overlap
Christopher Thai1,2, Amartya Singh1,2, Daniel Herranz3,4,5
1Rutgers Cancer Institute, Rutgers University, New Brunswick, NJ, 08901, USA.
Abstract:
Single-cell RNA sequencing allows defining cellular identities based on transcriptional similarity using unsupervised clustering. However, a single clustering resolution may not yield groups of cells that represent both broad, well-defined populations and smaller subpopulations simultaneously. Therefore, when cell identities are not known prior to sequencing, robust comparison and annotation of inferred de novo clusters remains a challenge. Here, we introduce CANTAO, in which we propose the average overlap metric to define the distance between single-cell clusters by comparing ranked lists of differentially expressed genes in a top-weighted manner. We benchmark CANTAO in truth-known datasets comprised of similar yet distinct cell populations and show that evaluating clusters with average overlap results in a consistent, precise, and biologically meaningful recapitulation of true cell identities. We then analyze unsorted mouse thymocytes and characterize stages of T-cell development in the thymus, including minor populations of double-negative (CD4-CD8-) T cells that are difficult to confidently detect among unsorted single cells. We demonstrate that CANTAO enables robust, reproducible characterization of single-cell data and clarifies biological interpretation of underlying identities in homogeneous populations.
More Related Videos
Related Concept Videos
RNA-seq
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Genome Annotation and Assembly

