Related Experiment Video
Updated: Jul 15, 2026

A Method for 3D Reconstruction and Virtual Reality Analysis of Glial and Neuronal Cells
Published on: September 28, 2019
scDifformer: diffusion-based post-training for virtual cell modeling across large-scale single-cell data
Zhan Xiao1, Wuke Wang2, Xin Long1,3
1Research Center for Life Sciences Computing, Zhejiang Lab, Hangzhou, Zhejiang 311121, China.
Abstract:
Virtual cells represent a promising paradigm to understand cellular mechanisms, behavior, and dynamics. The realization of virtual cells relies on the accurate modeling of cellular dynamics from large-scale, multi-modal single-cell data. However, experiment-specific technical noise and intrinsic biological heterogeneity pose major challenges for virtual cell modeling. To address this gap, we present scDifformer, a context-aware transformer model augmented with a denoising diffusion module and a dedicated post-training phase. This three-phase design, comprising masked language model pre-training, diffusion-driven post-training, and downstream fine-tuning, directly enhances scDifformer's ability to denoise sparse, noisy data and generalize across studies. Benchmarking across seven tissues and multiple independent studies shows that the diffusion module consistently improves cross-dataset performance, particularly in settings with strong batch effects. By combining the strengths of transformer and diffusion models, scDifformer achieves state-of-the-art performance in cell type annotation across diverse datasets. It further demonstrates robust capability in resolving immune cell identities across multiple tissues, accurately recovering key marker genes, functional pathways, and cross-tissue differentiation trajectories. Finally, by integrating scDifformer with a graph neural network, we extend its utility to spatial transcriptomics, significantly enhancing spot-level deconvolution accuracy. Altogether, scDifformer provides a scalable and biologically grounded framework for modeling heterogeneous single-cell data, offering a powerful foundation for the development of high-fidelity, multi-modal virtual cell models.

