Related Experiment Video
Updated: Aug 12, 2026

A Versatile Automated Platform for Micro-scale Cell Stimulation Experiments
Published on: August 6, 2013
SKIM: A fast sketching strategy integrated with model's dynamic-feedback for large-scale single-cell transcriptomic
Jiaxing Bai1, Feng Zhou2, Chongyang Tan2
1Department of Automation, National Institute for Data Science in Health and Medicine, State Key Laboratory of Mariculture Breeding, Xiamen Key Laboratory of Big Data Intelligent Analysis and Decision, Xiamen University, Xiamen, Fujian, China.
Abstract:
The rapid expansion of single-cell RNA sequencing (scRNA-seq) datasets poses analytical challenges, including computational burden raised by increasing data scale, and imbalances in cell population sizes. Dataset sketching mitigates these issues by selecting representative subsets that preserve biological signals. Existing sketching methods rely on geometric distances between transcriptomic expression profiles to select representative cells that cover the expression space. While effective at capturing global expression patterns, these methods are insensitive to local expression variations that encode multiscale transcriptomic differences. Moreover, by prioritizing maximal coverage of the expression space, these methods overselect peripheral cells, including rare and abnormal cells. To overcome these limitations, we propose SKIM, a fast SKetching strategy integrated with model's dynamic-feedback for large-scale single-cell transcriptomic analysis. Dynamic-feedback is defined as the sequence of reconstruction losses for each cell across training epochs, reflecting how the model progressively learns the transcriptomic expression profile. During training, as the model captures features from global expression patterns to local variations, dynamic-feedback from loss trajectories provides comprehensive multiscale views of the variation between cells. Based on dynamic-feedback, SKIM identifies and removes abnormal cells exhibiting unstable feedback patterns, and constructs sketches through clustering and size-aware sampling in the dynamic-feedback space, reducing data size while balancing cell populations. Across nine benchmark datasets, SKIM outperforms four state-of-the-art sketching methods in cell type annotation, data integration, bulk RNA-seq deconvolution, and developmental trajectory inference. Moreover, SKIM demonstrates high computational efficiency, achieving over a 20-fold speedup compared with existing methods when generating a 10% sketch from 216,611 cells. Overall, SKIM provides an efficient framework for large-scale dataset sketching that effectively retains critical biological signals while reducing computational burden.

