MegSite:基于多模式蛋白语言模型的准确核酸结合残留预测方法
Feng Hu1, Wenwu Zeng2, Shaoliang Peng2
1College of Computer and Artificial Intelligence, Zhengzhou University, Zhengzhou 450001, China.
Briefings in bioinformatics
|October 4, 2025
概括
MegSite使用多模式蛋白质特征从序列,结构和功能准确地识别核酸结合部位. 这种新的方法超越了现有的工具,推进了基因表达和调控机制研究.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 结构生物学 结构生物学
背景情况:
- 准确识别核酸结合残留物对于理解基因表达和调节机制至关重要.
- 由于蛋白质数据的复杂性,当前的计算方法难以高准确度.
研究的目的:
- 介绍MegSite,一个新的多模式蛋白质语言模型信息方法,用于核酸结合部位的预测.
- 整合各种蛋白质特征,包括序列,结构和功能,以提高预测准确度.
主要方法:
- 开发了MegSite,一种利用多式蛋白质语言模型特征的方法,特别是ESM3.
- 综合区分知识从蛋白质序列,结构和功能.
- 在多个独立测试集 (DNA-129_Test,DNA-181_Test,RNA-117_Test,RNA-285_Test) 上评估的性能.
主要成果:
- MegSite显著优于现有的核酸结合部位预测方法.
- 与第二个最佳方法相比,在所有测试的数据集上实现了改善的马修斯相关系数.
- 在结构相似性较低的蛋白质上表现强,优于基于结构的方法.
结论:
- MegSite的卓越性能源于其有效地整合了多模式蛋白质知识.
- 该方法显示了广泛的适用性,扩展到预测的蛋白质结构和新的RNA结合残留数据集.
- MegSite在预测核酸结合部位方面取得了重大进展.
相关概念视频
Conserved Binding Sites
5.0K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
5.0K
Ligand Binding Sites
14.9K
Proteins are dynamic macromolecules that carry out a wide variety of essential processes; however, the activities of most proteins depend on their interactions with other molecules or ions, known as ligands.
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
14.9K
Protein-protein Interfaces
14.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.4K
Protein Networks
4.5K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.5K
Leaky Scanning
5.6K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.6K
Improving Translational Accuracy
14.1K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
14.1K


