研究同进化信号的统计条件,使蛋白质伙伴的算法预测成为可能
José Fiorote1, João Alves1, Letícia Stock2
1Laboratório de Biologia Teórica e Computacional (LBTC), Universidade de Brasília, Brasilia, DF 70910-900, Brasil.
Journal of chemical information and modeling
|April 15, 2025
概括
这项研究引入了马尔科夫模型,用于使用共同进化数据预测蛋白质伙伴. 它发现,忽视微小的序列差异可以提高大蛋白家族的预测准确性.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 统计建模 统计建模
背景情况:
- 预测蛋白质与蛋白质相互作用对于理解生物系统至关重要.
- 氨基酸序列中的共同进化信号提供了一条推断蛋白质伙伴关系的途径.
- 目前的方法面临着大型蛋白质家族和序列相似性的挑战.
研究的目的:
- 从序列数据中研究能够对蛋白质伙伴进行算法预测的统计条件.
- 根据共同进化信息开发一个预测模型.
- 确定共进化预测模型的局限性和潜在改进.
主要方法:
- 马尔科夫随机模型的开发.
- 用正常分布的波桑混合来计算状态概率.
- 关键参数分析:总序列 (M),同进化差距 (α) 和方差 (σ02).
主要成果:
- 算法方法最大化同进化信息与大型蛋白质家族 (M ≥100) 的斗争.
- 真正阳性率通过不考虑类似序列之间的不匹配而增加.
- 基于{α, σ02},可以区分优化和退化的解决方案.
结论:
- 该研究为分类可靠的合作伙伴预测的蛋白质家族提供了一个框架.
- 忽视类似序列之间的微不足道错误可以提高预测的准确性.
- 促进了对大规模蛋白质数据分析的共同进化模型的理解.
更多相关视频
相关概念视频
Protein-protein Interfaces
12.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.4K
Conserved Binding Sites
4.1K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.1K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Conservation of Protein Domains Over Different Proteins
10.6K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.6K
Signal Sequences and Sorting Receptors
5.1K
Signal sequences are short amino acid sequences that guide newly synthesized proteins to their proper location within the cell. Classical signal sequences are fifteen to sixty amino acids long and present at the N-terminus of a polypeptide chain. Each signal sequence has a conserved segment of basic residues towards their N terminus, a hydrophobic core, and a C-terminus rich in polar residues. The C-terminus also contains a signal cleavage site and features a -3 -1 sequence motif. The -3-1...
5.1K
Evolutionary Relationships through Genome Comparisons
5.6K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.6K


