The impacts of active and self-supervised learning on efficient annotation of single-cell expression data
Michael J Geuenich1,2, Dae-Won Gong3, Kieran R Campbell4,5,6,7,8,9
1Lunenfeld-Tanenbaum Research Institute, Sinai Health System, Toronto, ON, M5G 1×5, Canada. mgeuenich@lunenfeld.ca.
Nature Communications
|February 2, 2024
Summary
Active and self-supervised learning reduce cell annotation time and cost in single-cell analysis. This study benchmarks these methods, introducing adaptive reweighting for improved accuracy, especially with marker knowledge.
Area of Science:
- Computational Biology
- Genomics
- Machine Learning
Background:
- Single-cell data analysis requires accurate cell type and state annotation.
- Manual cell labeling for training datasets is labor-intensive and costly.
- Active and self-supervised learning offer potential solutions to reduce annotation burden.
Purpose of the Study:
- To comprehensively benchmark active and self-supervised learning strategies for single-cell annotation.
- To evaluate the performance of these methods across different single-cell technologies and annotation algorithms.
- To assess the impact of cell type imbalance and similarity on annotation strategies.
Main Methods:
- Benchmarking of active learning and self-supervised labeling strategies.
- Evaluation across diverse single-cell technologies and cell type annotation algorithms.
- Introduction of adaptive reweighting, including a marker-aware version.
Main Results:
- Active and self-supervised strategies offer significant benefits for single-cell annotation.
- Adaptive reweighting demonstrates competitive performance compared to existing methods.
- Prior knowledge of cell type markers substantially improves annotation accuracy.
Conclusions:
- Active and self-supervised learning are effective for efficient single-cell annotation.
- The proposed adaptive reweighting method provides a valuable tool for researchers.
- Leveraging marker information enhances the precision of automated cell type identification.


