Related Experiment Video
Updated: Jun 17, 2026

A Virtual Machine Platform for Non-Computer Professionals for Using Deep Learning to Classify Biological Sequences of Metagenomic Data
Published on: September 25, 2021
Language model-based self-training reduces labeled data requirements by 99% for biological sequence classification
Jingwen Liu1, Danmo Gao1, Yan Yuan2
1School of Computer Science and Artificial Intelligence, Hubei University of Technology, 28 Nanli Road, Hongshan District, Wuhan 430068, China.
This study introduces a novel framework integrating pre-trained language models (PLMs) with semi-supervised learning (SSL) for biological sequence function prediction. The method significantly enhances accuracy with minimal labeled data, outperforming traditional approaches.
Area of Science:
- Computational Biology
- Bioinformatics
- Genomics
Background:
- Predicting biological sequence function is crucial for understanding disease mechanisms and genetic variation.
- Existing methods face challenges due to limited labeled data and complex sequence context modeling.
- Previous research has explored semi-supervised learning (SSL) and pre-trained language models (PLMs) separately, overlooking their combined potential.
Purpose of the Study:
- To develop an integrated framework combining PLMs and SSL for improved biological sequence function prediction.
- To leverage PLMs for feature extraction and SSL for decision boundary refinement.
- To demonstrate the framework's effectiveness in low-resource settings for tasks like DNA-binding protein (DBP) and non-coding RNA (ncRNA) detection.
Main Methods:
- Utilized PLMs as feature extractors to capture sequence semantics from large unlabeled datasets.
- Employed SSL with confidence-weighted pseudo-label selection to constrain the model's decision boundary.
- Applied the integrated framework to DNA-binding protein (DBP) and non-coding RNA (ncRNA) prediction tasks.
Main Results:
- Achieved competitive performance compared to fully supervised methods using significantly fewer labeled samples (as little as 1%).
- Demonstrated superior performance over traditional SSL methods like TSVM through a language model-based self-training approach.
- Successfully identified novel biomolecules, highlighting the framework's efficacy in low-resource scenarios.
Conclusions:
- The proposed framework offers an efficient solution for biological sequence classification, particularly in data-scarce environments.
- Integrating PLMs and SSL provides a powerful methodology for deciphering biological sequence function.
- This approach lays a foundation for advancing computational biology and discovering new biomolecules.
Related Concept Videos
Evolutionary Relationships through Genome Comparisons
Improving Translational Accuracy
Improving Translational Accuracy
Per-Unit Sequence Models
Zero-sequence currents, which are identical in magnitude and phase, generate a neutral current, resulting in voltage drops across the neutral impedance and the low-voltage winding. If the...