An overview of self-supervised deep learning applications to molecular data
Lluis Borràs Ferrís1,2, Riccardo Fratti1, Valerio Nucera1
1Institute of Informatics, University of Applied Sciences Western Switzerland (HES-SO), Technopôle 3, 3960 Sierre, Valais, Switzerland.
Briefings in Bioinformatics
|July 17, 2026
Summary
Self-supervised learning (SSL) offers a powerful way to analyze vast unlabeled molecular data for bioinformatics. This review highlights SSL
Area of Science:
- Bioinformatics
- Computational Biology
- Machine Learning
Background:
- High-throughput sequencing generates massive molecular data, challenging traditional analysis.
- Deep learning (DL) excels but often requires large labeled datasets.
- Self-supervised learning (SSL) leverages unlabeled data for representation learning.
Purpose of the Study:
- To provide a comprehensive review of self-supervised learning applications in omics data analysis.
- To guide researchers in deep learning and bioinformatics on utilizing SSL for molecular data.
- To foster future research at the intersection of SSL and omics.
Main Methods:
- Review of 17 studies applying SSL to various omics data types.
- Categorization of applications by data type, common tasks, and model architectures.
- Examination of SSL principles, including foundation models.
Main Results:
- SSL is effectively applied to molecular sequences for learning meaningful representations.
- Key applications like DNABERT and Nucleotide Transformer demonstrate SSL's utility in gene regulation studies.
- Identified common tasks, model architectures, and relevant repositories for SSL in omics.
Conclusions:
- SSL presents a promising approach for analyzing large-scale unlabeled omics data.
- Future directions include multi-omics data integration and advanced pretext task development.
- SSL can significantly advance our understanding of molecular mechanisms and gene regulation.

