LMCrot:通过利用基于变压器的蛋白质语言模型的可解释窗口级嵌入来进行增强的蛋白质crotonylation位点预测器
Pawel Pratyush1, Soufia Bahmani1, Suresh Pokharel1
1Department of Computer Science, Michigan Technological University, Houghton, MI 49931, United States.
Bioinformatics (Oxford, England)
|April 25, 2024
概括
这项研究探讨了蛋白质语言模型,用于预测蛋白质化部位. 一种新型模型,LMCrot,通过整合蛋白序列和属性信息来提高预测准确性,优于现有方法.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 机器学习 机器学习
背景情况:
- 蛋白质语言模型 (pLMs) 为下游任务提供了强大的全球上下文化表示.
- 使用pLMs在每残留水平上预测像crotonylation (Kcr) 这样的翻译后修改是一个新兴的领域.
- 需要有效的策略来在pLM中编码兴趣点信息,以进行每余量预测.
研究的目的:
- 调查和优化使用pLMs进行每残留克罗化 (Kcr) 预测.
- 开发新的深度学习架构,利用pLM嵌入式进行改进的Kcr预测.
- 通过融合策略提高模型解释性和预测性能.
主要方法:
- 实验了不同的输入序列类型 (全长与窗口) 和plm的嵌入策略.
- 开发了T5ResConvBiLSTM,一个剩余的ConvBiLSTM网络处理ProtT5-XL-UniRef50窗口嵌入.
- 使用注意力权重和SHAP值来解释ProtT5嵌入.
- 实施了一种堆叠的泛化方法 (LMCrot),将ProtT5嵌入物与本地氨基酸属性和监督嵌入物融合在一起.
主要成果:
- 在三个数据集中,T5ResConvBiLSTM模型超过了最先进的Kcr预测器.
- 模型解释性分析验证了基于完整序列的窗口嵌入的实用性.
- 集成本地特征的LMCrot模型进一步显著改善了跨数据集的预测性能.
结论:
- 来自pLM的嵌入,特别是来自完整序列的窗口级嵌入,对于像Kcr.这样的每残余预测任务是有效的.
- 新的深度学习架构和功能融合策略可以大幅提高Kcr预测的准确性.
- 开发的LMCrot模型为Kcr站点预测提供了一个强大而可解释的工具.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Insertion of Single-pass Transmembrane Proteins in the RER
6.7K
Integral membrane proteins are proteins adhered to the lipid bilayer of a cell organelle or membrane. They can be of two types: transmembrane integral proteins that span the lipid bilayer and monotopic proteins that are attached to either side of the membrane but do not pass through it.
Integral transmembrane proteins possess transmembrane and extra membrane domains. The transmembrane domains are primarily made of 20-25 hydrophobic amino acids arranged in a helical secondary confirmation. These...
Integral transmembrane proteins possess transmembrane and extra membrane domains. The transmembrane domains are primarily made of 20-25 hydrophobic amino acids arranged in a helical secondary confirmation. These...
6.7K
Cotranslational Protein Translocation
7.3K
Translocation of proteins across membranes is an ancient process that occurs even in bacteria and archaebacteria. In fact, the components of the translocation machinery are still conserved between prokaryotes and eukaryotes.
Sec61 channel partners for cotranslational translocation
During cotranslational translocation, the Sec61 channel partners with the signal recognition particle (SRP), the signal recognition particle receptor (SR), and the ribosomes to transport the nascent polypeptide chain...
Sec61 channel partners for cotranslational translocation
During cotranslational translocation, the Sec61 channel partners with the signal recognition particle (SRP), the signal recognition particle receptor (SR), and the ribosomes to transport the nascent polypeptide chain...
7.3K
Improving Translational Accuracy
10.2K
Base complementarity between the three base pairs of mRNA codon and the tRNA anticodon is not a failsafe mechanism. Inaccuracies can range from a single mismatch to no correct base pairing at all. The free energy difference between the correct and nearly correct base pairs can be as small as 3 kcal/ mol. With complementarity being the only proofreading step, the estimated error frequency would be one wrong amino acid in every 100 amino acids incorporated. However, error frequencies observed in...
10.2K
Leaky Scanning
5.1K
During most eukaryotic translation processes, the small 40S ribosome subunit scans an mRNA from its 5' end until it encounters the first start AUG codon. The large 60S ribosomal subunit then joins the smaller one to initiate protein synthesis. The location of the translation initiation is largely determined by the nucleotides near the start codon as there may be multiple translation initiation sites present on the mRNA. Marilyn Kozak discovered that the sequence RCCAUGG (where R...
5.1K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K


