通过蛋白质语言模型检测循环顺序
Yue Hu1,2, Bin Huang3, Chun Zi Zang2
1School of Bioengineering, Qilu University of Technology (Shandong Academy of Sciences), Jinan, Shandong 250300, China.
Computational and structural biotechnology journal
|January 27, 2025
概括
检测蛋白质循环排列 (CP) 是至关重要的. 一种新的方法,plmCP,使用遗传原理和蛋白质语言模型来克服传统序列和基于结构的方法的局限性,以准确识别CP.
科学领域:
- 生物化学 生物化学
- 计算生物学 计算生物学
- 结构生物学 结构生物学
背景情况:
- 蛋白质循环合 (CP) 对于理解蛋白质的演变和功能很重要.
- 现有的检测方法,包括基于序列和基于结构的方法,在准确性和范围上都有局限性.
- 蛋白质语言模型 (PLM) 显示出基于序列的分析的希望,但由于线性序列依赖性,与CP检测扎.
研究的目的:
- 开发一种新的方法来准确检测蛋白质的循环变换.
- 克服现有的序列对齐算法的局限性,以识别CP.
- 为了利用PLM的力量与基因原理相结合,改善CP检测.
主要方法:
- 开发了plmCP,这是一种新的方法,将经典遗传原理与基于PLM的先进对齐技术相结合.
- 该方法旨在规避传统对齐算法固有的线性序列顺序依赖.
- 采用PLM来有效地利用序列信息,而不需要结构输入.
主要成果:
- plmCP成功地解决了当前基于PLM的对齐工具在检测循环排列时所面临的挑战.
- 整合遗传知识使该方法能够克服序列顺序的依赖性.
- 证明有效检测循环变换,增强蛋白质研究能力.
结论:
- plmCP在检测蛋白质循环变异方面取得了重大进展.
- 该方法能够处理结构灵活性并避免依赖序列顺序的能力扩大了其适用性.
- 通过提供强大的CP识别工具,为增强蛋白质研究和工程做出贡献.
相关概念视频
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Signal Sequences and Sorting Receptors
5.2K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.2K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Improving Translational Accuracy
8.6K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
8.6K
Proteomics
7.2K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.2K
Protein Complexes with Interchangeable Parts
2.5K
Groups of proteins may form a complex where each protein in this complex has a different role in the overall execution of the complex’s function. Often some of the proteins in the complex can be replaced by a closely related variant to give a complex that contains many of the same components yet is functionally distinct.
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
The SCF ubiquitin ligase is a protein complex of five individual proteins. This complex attaches ubiquitin to other target proteins to mark them for degradation. In order...
2.5K


