Related Experiment Video
Updated: May 15, 2025

13:42
RNA Secondary Structure Prediction Using High-throughput SHAPE
Published on: May 31, 2013
31.3K
Comprehensive benchmarking of large language models for RNA secondary structure prediction
Luciano I Zablocki1, Leandro A Bugnon1, Matias Gerard1
1Research Institute for Signals, Systems and Computational Intelligence, sinc (i), FICH-UNL/CONICET, Ruta Nacional Nº 168, km 472.4, Santa Fe (3000), Argentina.
Briefings in Bioinformatics
|April 10, 2025
Summary
Large language models (LLMs) for RNA show promise for predicting RNA secondary structures. However, their generalization capabilities vary, with two models outperforming others, especially in low-homology scenarios.
Area of Science:
- Computational Biology
- Bioinformatics
- Machine Learning
Background:
- Large language models (LLMs) have shown success in modeling DNA and protein sequences.
- LLMs for RNA are emerging, learning rich numerical representations of RNA bases from massive datasets.
- High-quality RNA representations are hypothesized to improve performance on data-intensive tasks like RNA secondary structure prediction.
Purpose of the Study:
- To comprehensively analyze and compare pretrained RNA LLMs for RNA secondary structure prediction.
- To evaluate the generalization capabilities of RNA LLMs on new and diverse RNA structures.
- To establish a unified experimental setup and curated benchmark datasets for RNA LLM evaluation.
Main Methods:
- Utilized a common deep learning architecture to assess RNA LLM representations for secondary structure prediction.
- Evaluated multiple pretrained RNA LLMs on benchmark datasets with increasing generalization difficulty.
- Conducted a comparative analysis of RNA LLM performance across different homology levels.
Main Results:
- Two specific RNA LLMs demonstrated superior performance compared to other models.
- Significant challenges were identified in the generalization capabilities of RNA LLMs, particularly in low-homology scenarios.
- The study provides curated benchmark datasets and a unified experimental framework.
Conclusions:
- Pretrained RNA LLMs can enhance RNA secondary structure prediction, but their generalization abilities require further investigation.
- Model selection is critical, as performance varies significantly among different RNA LLMs.
- Future research should focus on improving RNA LLM generalization, especially for sequences with low homology to training data.
Related Concept Videos
RNA-seq
9.7K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
9.7K
Ribosome Profiling
3.4K
Ribosome profiling or ribo-sequencing is a deep sequencing technique that produces a snapshot of active translation in a cell. It selectively sequences the mRNAs protected by ribosomes to get an insight into a cell’s translation landscape at any given point in time.
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
Applications of ribosome profiling
Ribosome profiling has many applications, including in vivo monitoring of translation inside a particular organ or tissue type and quantifying new protein synthesis levels.
The technique...
3.4K
Leaky Scanning
5.0K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.0K
Improving Translational Accuracy
8.5K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.5K
RNA Structure
4.5K
The basic structure of RNA consists of a string of ribonucleotides attached by phosphodiester bonds. Although most RNA is single-stranded, it can form complex secondary and tertiary structures. Such structures play essential roles in the regulation of transcription and translation.
Different Types of RNA Have the Same Basic Structure
There are three main types of ribonucleic acid (RNA) involved in protein synthesis: messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA). All three...
Different Types of RNA Have the Same Basic Structure
There are three main types of ribonucleic acid (RNA) involved in protein synthesis: messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA). All three...
4.5K
RNA Stability
33.1K
Intact DNA strands can be found in fossils, while scientists sometimes struggle to keep RNA intact under laboratory conditions. The structural variations between RNA and DNA underlie the differences in their stability and longevity. Because DNA is double-stranded, it is inherently more stable. The single-stranded structure of RNA is less stable but also more flexible and can form weak internal bonds. Additionally, most RNAs in the cell are relatively short, while DNA can be up to 250 million...
33.1K

