Related Experiment Video
Updated: Jul 1, 2026

Droplet Barcoding-Based Single Cell Transcriptomics of Adult Mammalian Tissues
Published on: January 10, 2019
scCCVGBen for benchmarking of single-cell representation learning anchored on a centroid-coupled variational graph
Zeyu Fu1, Jiawei Fu2, Chunlin Chen3
1State Key Laboratory of Trauma and Chemical Poisoning, Institute of Combined Injury, Chongqing Engineering Research Center for Nanomedicine, College of Preventive Medicine, Army Medical University, Chongqing, China.
None:
Single-cell omics routinely profile millions of cells across the transcriptome and the epigenome. However, embeddings used for clustering, trajectory inference, and visualization remain unstable: stochastic variational autoencoders inject sampling noise at inference, and methods reported on idiosyncratic cohorts defeat head-to-head comparison. We introduce scCCVGBen, a benchmark of single-cell representation-learning methods. Its reference configuration is a centroid-coupled variational graph autoencoder built from three design choices: the centroid (deterministic posterior mean) used as the inference embedding, a coupling-regularized dual-reconstruction bottleneck, and a graph attention encoder over a -nearest-neighbor cell-cell graph. We assess this configuration within a decoupled benchmark that varies the algorithmic core, encoder backbone, graph construction, dataset cohort, and evaluation suite as independent axes. The cohort, drawn from the Gene Expression Omnibus (GEO) and the European Nucleotide Archive (ENA), balances scRNA-seq and scATAC-seq equally and spans hematopoiesis, neuronal differentiation, immune populations, organ atlases, tumor microenvironments, and developmental time courses. Across the cohort, scCCVGBen improves average silhouette width by and intrinsic-overall geometry by over a stochastic variational encoder (VAE) on paired scRNA-seq; gains over scVI reach and , and on scATAC-seq, the gain over PeakVI on intrinsic geometry reaches . Robustness analyses across 14 graph encoders and 5 graph-construction strategies show where alternative architectures remain competitive. Three paired hematopoietic case studies: sleep-disrupted bone marrow alongside a gastric tumor atlas, cord blood megakaryopoiesis alongside aged hematopoietic stem cells, and radiation-injury hematopoiesis alongside the COVID-19 bronchoalveolar landscape, recover coherent latent-gene programs spanning hematopoietic, epithelial-stromal, megakaryocytic, and antiviral-macrophage axes. The benchmark cohort, per-method scores, and per-dataset metadata are released through three companion sites: a Hugo atlas, a Next.js interactive cohort browser, and a cross-tool discovery surface, so the cohort can be inspected without cloning the source repository. The result is a stable, interpretable embedding that carries cleanly from benchmarking to biological discovery.
