A benchmark of batch-effect correction methods for single-cell RNA sequencing data
Hoa Thi Nhu Tran1, Kok Siong Ang1, Marion Chevrier1
1Singapore Immunology Network (SIgN), Agency for Science, Technology and Research (A*STAR), 8A Biomedical Grove, Immunos Building, Level 3, Singapore, 138648, Singapore.
Genome Biology
|January 18, 2020
Summary
This study benchmarks 14 batch correction methods for single-cell RNA sequencing (scRNA-seq) data. Harmony, LIGER, and Seurat 3 are recommended for effective batch integration, with Harmony being the fastest option.
Area of Science:
- Computational biology
- Genomics
- Bioinformatics
Background:
- Single-cell RNA sequencing (scRNA-seq) data present batch effects from different technologies, challenging data integration.
- Effective batch correction is crucial for analyzing the growing volume of scRNA-seq data.
Purpose of the Study:
- To benchmark and compare available computational methods for batch-effect removal in scRNA-seq data.
- To identify the most suitable methods for accurate and efficient data integration.
Main Methods:
- Evaluated 14 batch correction algorithms on scRNA-seq datasets.
- Assessed methods based on computational runtime, scalability, and batch-effect removal efficacy.
- Utilized metrics like kBET, LISI, ASW, and ARI across five distinct experimental scenarios.
- Investigated the impact of batch correction on differential gene expression analysis.
Main Results:
- Compared 14 batch correction methods across various scenarios including identical/non-identical cell types, multiple batches, large datasets, and simulated data.
- Performance was evaluated using kBET, LISI, ASW, and ARI metrics.
- Harmony, LIGER, and Seurat 3 demonstrated strong performance in batch integration and cell type purity preservation.
Conclusions:
- Harmony, LIGER, and Seurat 3 are recommended for scRNA-seq data integration.
- Harmony is the preferred initial choice due to its superior computational efficiency.
- The study provides guidance for selecting appropriate batch correction tools for large-scale transcriptomic data analysis.


