Related Experiment Video
Updated: Jun 8, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
505
Semi-supervised learning with pseudo-labeling compares favorably with large language models for regulatory sequence
Han Phan1, Céline Brouard1, Raphaël Mourad1,2
1INRAE, MIAT, 31326 Castanet-Tolosan, France.
Briefings in Bioinformatics
|November 3, 2024
Summary
This study introduces a semi-supervised learning method for predicting molecular processes from DNA sequences, leveraging unlabeled genomic data to improve deep learning models for non-coding variants.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Deep learning offers biological insights for non-coding single nucleotide polymorphisms (SNPs) from genome-wide association studies (GWAS).
- Supervised learning methods are limited by the scarcity of functional data for DNA sequences.
- Mammalian DNA sequence data is abundant but often lacks associated functional information.
Purpose of the Study:
- To develop a novel semi-supervised learning (SSL) approach to overcome limitations of supervised learning in predicting molecular processes from DNA.
- To enhance model pre-training by effectively utilizing large amounts of unlabeled DNA sequences.
- To improve predictive performance, especially in scenarios with limited functional data, such as predicting transcription factor binding.
Main Methods:
- Proposed a novel semi-supervised learning (SSL) method based on pseudo-labeling for pre-training models on unlabeled DNA sequences.
- Incorporated principles from the Noisy Student algorithm to refine pseudo-labeled data confidence during pre-training.
- The flexible approach allows training of various neural network architectures, including advanced models.
Main Results:
- The SSL approach demonstrated significant improvements in predictive performance compared to standard supervised learning across various tasks.
- The method showed particular effectiveness for predicting transcription factor binding sites, even with very limited training data.
- Small models trained using SSL achieved performance comparable to or exceeding that of large language models like DNABERT2.
Conclusions:
- Semi-supervised learning, particularly with pseudo-labeling and Noisy Student principles, is a powerful strategy for deep learning in genomics.
- This approach effectively leverages abundant unlabeled DNA sequences to enhance predictive modeling for molecular processes.
- SSL provides a flexible and high-performing alternative to supervised learning, especially when functional genomic data is scarce.
Related Concept Videos
Cis-regulatory Sequences
9.8K
Cis-regulatory sequences are short fragments of non-coding DNA that are present on the same chromosomes as the genes that they regulate. These fragments serve as binding sites for transcriptional regulators, proteins that are responsible for controlling gene transcription and differential gene expression across cell types in eukaryotes. Cis-regulatory sequences can be close to the gene of interest or thousands of bases away in the DNA sequence; however, those sequences that are further away are...
9.8K
Cooperative Binding of Transcription Regulators
6.4K
Transcriptional regulators bind to specific cis-regulatory sequences in the DNA to regulate gene transcription. These cis-regulatory sequences are very short, usually less than ten nucleotide pairs in length. The short length means that there is a high probability of the exact same sequence randomly occurring throughout the genome. Since regulators can also bind to groups of similar sequences, this further increases the chances of random binding. Transcriptional regulators form...
6.4K
Regulation of Expression Occurs at Multiple Steps
3.0K
3.0K
Master Transcription Regulators
6.9K
Master transcription regulators are regulatory proteins that are predominantly responsible for regulating the expression of multiple genes. Often these genes work in concert to drive a complex process. Activation of a master transcription regulator can lead to a cascade of transcriptional activation necessary for that outcome. These regulators can directly bind to the regulatory sequences of the various genes involved, or they can indirectly regulate transcription by binding to regulatory...
6.9K
Regulation of Expression at Multiple Steps
873
The gene expression in cells is regulated at different stages: (i) transcription, (ii) RNA processing, (iii) RNA localization, and (iv) translation. Transcriptional regulation is mediated by regulatory proteins such as transcription factors, activators, or repressors—these control gene expression by initiating or inhibiting the transcription of genes. Once a precursor or pre-mRNA is produced, it undergoes post-transcriptional modification, including 5' capping, splicing, and the...
873
lncRNA - Long Non-coding RNAs
8.5K
In humans, more than 80% of the genome gets transcribed. However, only around 2% of the genome codes for proteins. The remaining part produces non-coding RNAs which includes ribosomal RNAs, transfer RNAs, telomerase RNAs, and regulatory RNAs, among other types. A large number of regulatory non-coding RNAs have been classified into two groups depending upon their length – small non-coding RNAs, such as microRNA, which are less than 200 nucleotides in length, and long non-coding RNA...
8.5K

