Related Experiment Video
Updated: Sep 25, 2026

Characterizing Mutational Load and Clonal Composition of Human Blood
Published on: July 11, 2019
Cell-level random splits leak group-owned answers in single-cell benchmarks
Abstract:
Machine learning models in single-cell biology increasingly forecast differentiation, reprogramming and therapeutic response from early transcriptomic profiles. Testing whether a model has learned real biology requires held-out cells. Single-cell data, however, are grouped: cells from the same clone, patient or batch share the same label. A random split therefore places relatives of each test cell, carrying its label, in the training set, and a model can score well by memorizing a relative instead of learning a transferable rule. Grouped validation removes this leakage but leaves far fewer independent units behind each error bar. Here we show how to estimate this leakage before training any model, from two properties of the data: exposure, the fraction of test cells with relatives in training, and retrievability, how often a nearest-neighbor search returns such a relative rather than an unrelated cell. Across lineage-barcoded and patient data, exposure determines whether a random split opens a leakage channel, and retrievability determines how much it can inflate the score. The inflation is negligible where cell state has decoupled from ancestry, much larger where clonal sisters remain close in expression space, and in a patient cohort large enough to overturn a clinical conclusion. We also provide leakcheck , which computes both properties in seconds, before the outcome model is fitted.
More Related Videos
10:20Simultaneous Assessment of Kinship, Division Number, and Phenotype via Flow Cytometry for Hematopoietic Stem and Progenitor Cells
Published on: March 24, 2023
09:34A Combinatorial Single-cell Approach to Characterize the Molecular and Immunophenotypic Heterogeneity of Human Stem and Progenitor Populations
Published on: October 25, 2018