DiffPaSS-使用软分数的蛋白质序列的高性能可微分配对
Umberto Lupo1,2, Damiano Sgarbossa1,2, Martina Milighetti3,4
1Institute of Bioengineering, School of Life Sciences, École Polytechnique Fédérale de Lausanne (EPFL), Lausanne CH-1015, Switzerland.
Bioinformatics (Oxford, England)
|December 13, 2024
概括
我们开发了DiffPaSS,这是一种快速,灵活和无超参数的计算方法,用于识别相互作用的蛋白质序列. 该工具改进了现有的序列配对算法,并有助于预测蛋白质复杂结构.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 结构生物学 结构生物学
背景情况:
- 识别相互作用的蛋白质序列对于理解生物通路和功能至关重要.
- 相互作用的蛋白质在氨基酸使用中表现出进化相似性和共发生模式.
- 现有的方法通常依赖于对序列配对的近似优化,限制了灵活性和速度.
研究的目的:
- 引入DiffPaSS (使用软分数的可差配对),这是一个新的可差配对框架,用于优化相互作用的生物序列对.
- 提供灵活,快速和无超参数的解决方案,用于跨各种分数系统的序列配对.
- 证明DiffPaSS在预测蛋白质复杂结构和分析非对齐序列方面的实用性.
主要方法:
- 开发了DiffPaSS,这是一个在PyTorch中实现的可微分优化框架.
- 将DiffPaSS应用到一个 prokaryotic 数据集中,使用相互信息和邻近图形对齐得分.
- 对已有的序列配对算法进行DiffPaSS性能评估.
主要成果:
- 在优化序列相似性和共同进化得分方面,DiffPaSS显著优于现有的算法.
- 由DiffPaSS生成的配对对齐对于预测蛋白质复杂结构是有效的.
- DiffPaSS成功地处理非对齐的序列,如T细胞受体数据所示.
结论:
- DiffPaSS提供了一种卓越,高效和多用途的方法来识别相互作用的蛋白质序列.
- 该框架处理多种分数和非对齐序列的能力扩大了其在生物信息学中的适用性.
- DiffPaSS为推进蛋白相互作用和结构生物学研究提供了有价值的工具.
相关概念视频
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Conservation of Protein Domains
3.1K
3.1K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Wilcoxon Signed-Ranks Test for Matched Pairs
87
The Wilcoxon signed-rank test for matched pairs evaluates the null hypothesis by combining the ranks of differences with their signs. It essentially tests whether the median of the differences in a population of matched pairs is zero. Since the test incorporates more information than the sign test, it generally yields more trustable conclusions. This test also does not require the data to follow a normal distribution, but two conditions must be met for it to be applicable: (1) the data must...
87
Protein Complexes with Interchangeable Parts
1.8K
1.8K
Peptide Identification Using Tandem Mass Spectrometry
6.4K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.4K


