PairK:双向k-mer对齐用于量化无序区域中的蛋白质基因保护
Jackson C Halpin1, Amy E Keating1,2,3
1MIT Department of Biology, 77 Massachusetts Ave., Cambridge, MA 02139.
bioRxiv : the preprint server for biology
|August 2, 2024
概括
预测蛋白质相互作用至关重要. 一种名为PairK的新方法准确地量化了无序区域中的图案保护,改善了对生物相关蛋白质-蛋白质相互作用的识别.
科学领域:
- 分子生物学分子生物学
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
背景情况:
- 蛋白质与蛋白质之间的相互作用是细胞过程的基础.
- 这些相互作用往往通过在无序蛋白质区域内与短线性基因 (SLiM) 结合的域进行介导.
- 预测这些相互作用对于理解生物网络和制定假设至关重要,但目前的方法与错误阳性作斗争.
研究的目的:
- 开发一种新的计算方法,用于准确预测生物相关的短线性动图 (SLiM).
- 克服在无序蛋白区域内量化序列保存的局限性,这阻碍了SLiM识别.
- 改进功能域-SLiM相互作用的预测,用于绘制蛋白质网络.
主要方法:
- 介绍PairK (双向k-mer对齐),一种新的,无多重序列对齐 (MSA) 的方法.
- PairK 量化了特定于无序蛋白质区域内的图案保护.
- 与基于标准MSA的保护得分和基于大型语言模型 (LLM) 的预测器对比PairK的表现.
主要成果:
- 在识别生物学上重要的动机实例方面,PairK显著优于基于MSA和基于LLM的保护得分预测器.
- 该方法证明了在比传统的MSA更广泛的基因学距离上量化动机保护的能力.
- 研究结果表明,短线性图案 (SLiM) 可能比以前基于MSA的指标估计的更为保守.
结论:
- PairK提供了一种更有效的方法来量化无序区域中的图案保护,提高功能性蛋白质-蛋白质相互作用的预测.
- 这种方法解决了生物信息学的一个关键挑战,提高了SLiM识别的可靠性.
- 由于PairK的开源可用性,使其在生物学研究中更容易被采用,用于绘制交互网络和生成新假设.
相关概念视频
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Conservation of Protein Domains
3.1K
3.1K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Intrinsically Disordered Proteins
17.8K
Intrinsically disordered proteins are a group of proteins that do not fold into specific three-dimensional structures. Their structural flexibility allows them to complement ordered proteins to perform functions that are inaccessible to rigid structures. They are more common in eukaryotes than prokaryotes and may either be exclusively intrinsically disordered or hybrid proteins, consisting of a mix of ordered and disordered regions. The absence of a rigid structure in these proteins can be...
17.8K
Protein Organization
6.3K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.3K


