Fast analysis of scATAC-seq data using a predefined set of genomic regions
Valentina Giansanti1,2, Ming Tang3, Davide Cittaro2
1Department of Informatics, Systems and Communication, University of Milano-Bicocca, Milan, Italy.
F1000Research
|July 1, 2020
Summary
We present a faster method for single-cell ATAC sequencing (scATAC-seq) analysis using pseudoalignment. This approach reduces computational demands without sacrificing accuracy, making scATAC-seq data processing more accessible.
Area of Science:
- Genomics
- Computational Biology
- Bioinformatics
Background:
- Single-cell ATAC sequencing (scATAC-seq) analysis has scaled to thousands of cells.
- Current scATAC-seq pipelines require substantial computational resources.
- Alignment-free techniques have accelerated other single-cell data types.
Purpose of the Study:
- To introduce a pseudoalignment-based approach for scATAC-seq data analysis.
- To reduce computational time and hardware requirements for scATAC-seq.
- To maintain precision in scATAC-seq data processing.
Main Methods:
- Utilized public 10k PBMC and K562 cell line scATAC-seq data.
- Employed kallisto for pseudoalignment and bustools for quantification against DNase I Hypersensitive Sites (DHS) references.
- Compared results against the standard cellranger-atac pipeline.
Main Results:
- Kallisto-based quantification showed no bias for known peaks.
- Cell grouping and identification were consistent with standard methods.
- Analysis using DHS-derived references proved robust for cell identification and gene activity quantification.
Conclusions:
- The kallisto approach offers a significantly faster alternative for scATAC-seq analysis compared to standard pipelines.
- Using known DHS sites as a reference is effective for characterizing cell populations.
- This method enhances the efficiency of cell group labeling based on marker genes.


