多样化的数据库和机器学习模型来缩小RNA结构预测中的概括差距.
Albéric A de Lajarte1, Yves J Martin des Taillades2, Justin Aruda1
1Department of Microbiology, Harvard Medical School, Boston, MA, USA.
Science advances
|February 25, 2026
概括
研究人员开发了eFold,这是一种用于RNA二级结构预测的深度学习模型. 这种在各种化学探测数据上训练的模型通过结合结构复杂性,而不仅仅是数据库大小,提高了预测准确性.
科学领域:
- 分子生物学分子生物学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 大分子结构,特别是蛋白质和核酸,对于理解生物功能至关重要.
- 虽然蛋白质结构预测是先进的,但由于结构数据有限,RNA结构预测面临着挑战.
- 蛋白质数据库拥有超过18万个蛋白质结构,有助于深度学习的进步.
研究的目的:
- 为了解决RNA二次结构预测的局限性.
- 为微RNA和信使RNA区域提出新的二次结构模型.
- 开发和验证用于RNA结构预测的新深度学习架构.
主要方法:
- 使用化学探测生成了1098个初级微RNA和1456个人类信使RNA区域的二次结构模型.
- 开发了eFold,这是一个深度学习模型,灵感来自AlphaFold的Evoformer和传统架构.
- 在新生成的数据库和超过30万个现有的二次结构上训练了eFold.
主要成果:
- eFold在具有挑战性的,多样化的RNA结构测试集上展示了改进的预测性能.
- 新生成的数据集和eFold架构有助于提高预测准确性.
- 结果表明,整合结构多样性和复杂性是概括的关键.
结论:
- 仅扩大RNA结构数据库就不足以在不同家族中准确预测.
- eFold模型和多样化的数据集推进了RNA二次结构预测领域.
- 未来的努力应该集中在捕捉RNA结构的复杂性和多样性,以获得更好的模型.
相关概念视频
RNA-seq
12.2K
RNA sequencing, or RNA-Seq, is a high-throughput sequencing technology used to study the transcriptome of a cell. Transcriptomics helps to interpret the functional elements of a genome and identify the molecular constituents of an organism. Additionally, it also helps in understanding the development of an organism and the occurrence of diseases.
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
Before the discovery of RNA-seq, microarray-based methods and Sanger sequencing were used for transcriptome analysis. However, while...
12.2K
RNA Structure
79.5K
Overview
The basic structure of RNA consists of a five-carbon sugar and one of four nitrogenous bases. Although most RNA is single-stranded, it can form complex secondary and tertiary structures. Such structures play essential roles in the regulation of transcription and translation.
Different Types of RNA Have the Same Basic Structure
There are three main types of ribonucleic acid (RNA): messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA). All three RNA types consist of a...
The basic structure of RNA consists of a five-carbon sugar and one of four nitrogenous bases. Although most RNA is single-stranded, it can form complex secondary and tertiary structures. Such structures play essential roles in the regulation of transcription and translation.
Different Types of RNA Have the Same Basic Structure
There are three main types of ribonucleic acid (RNA): messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA). All three RNA types consist of a...
79.5K
RNA Structure
7.9K
The basic structure of RNA consists of a string of ribonucleotides attached by phosphodiester bonds. Although most RNA is single-stranded, it can form complex secondary and tertiary structures. Such structures play essential roles in the regulation of transcription and translation.
Different Types of RNA Have the Same Basic Structure
There are three main types of ribonucleic acid (RNA) involved in protein synthesis: messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA). All three...
Different Types of RNA Have the Same Basic Structure
There are three main types of ribonucleic acid (RNA) involved in protein synthesis: messenger RNA (mRNA), transfer RNA (tRNA), and ribosomal RNA (rRNA). All three...
7.9K
Nucleic Acid Structure
9.7K
The pentose sugar in DNA is deoxyribose, while in RNA the pentose sugar is ribose. The difference between the sugars is the presence of the hydroxyl group on the ribose's second carbon and a hydrogen on the deoxyribose's second carbon. The phosphate residue attaches to the hydroxyl group of the 5′ carbon of one sugar and the hydroxyl group of the 3′ carbon of the sugar of the next nucleotide, which forms a 5′ to 3′ phosphodiester linkage.
DNA Structure
DNA...
DNA Structure
DNA...
9.7K
Improving Translational Accuracy
15.3K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
15.3K
Improving Translational Accuracy
3.7K
3.7K


