Related Experiment Video
Updated: Jun 21, 2025

06:24
Multiplexed Analysis of Retinal Gene Expression and Chromatin Accessibility Using scRNA-Seq and scATAC-Seq
Published on: March 12, 2021
3.6K
Fast clustering and cell-type annotation of scATAC data using pre-trained embeddings
Nathan J LeRoy1,2, Jason P Smith1,3,4, Guangtao Zheng5
1Center for Public Health Genomics, School of Medicine, University of Virginia, Charlottesville, VA 22908, USA.
NAR Genomics and Bioinformatics
|July 8, 2024
Summary
This study introduces scEmbed, a novel machine learning framework for single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data. scEmbed leverages pre-trained models for efficient cell-type annotation and improved clustering.
Area of Science:
- Genomics
- Computational Biology
- Machine Learning
Background:
- Single-cell assay for transposase-accessible chromatin using sequencing (scATAC-seq) data is increasingly available.
- High dimensionality and sparsity in scATAC-seq data present significant computational challenges.
- Current methods generate cell embeddings in a single step, limiting flexibility.
Purpose of the Study:
- To develop a flexible and computationally efficient framework for analyzing scATAC-seq data.
- To introduce a transfer learning approach using pre-trained embedding models.
- To enable fast and accurate cell-type annotation without multi-modal data.
Main Methods:
- Implementation of scEmbed, an unsupervised machine learning framework.
- Learning low-dimensional embeddings of genomic regulatory regions.
- Utilizing pre-trained models on reference scATAC-seq data.
Main Results:
- scEmbed demonstrates strong clustering performance.
- The framework effectively learns transferable patterns of region co-occurrence.
- Pre-trained models facilitate rapid and accurate cell-type annotation.
Conclusions:
- Pre-trained embedding models offer a flexible and computationally advantageous approach for scATAC-seq analysis.
- scEmbed provides a robust tool for uncovering regulatory region co-occurrence patterns.
- The framework enables efficient cell-type annotation, reducing reliance on other data modalities.

