蛋白质家族,域和站点的PROSITE数据库
Christian J A Sigrist1, Béatrice A Cuche1, Edouard de Castro1
1Swiss-Prot Group, Swiss Institute of Bioinformatics (SIB), Centre Médical Universitaire (CMU), 1 rue Michel Servet, CH-1211 Geneva 4, Switzerland.
Nucleic acids research
|November 20, 2025
概括
蛋白质域数据库PROSITE现在集成AlphaFold结构并优化其搜索工具. 更新增强了SARS-CoV-2的研究,并揭示了新的转录因子链接.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 分子生物学分子生物学
背景情况:
- PROSITE是蛋白质域,家族和功能站点的关键数据库,使用模式和配置文件进行识别.
- ProRule通过关键的氨基酸信息增强了PROSITE的配置文件,有助于对UniProtKB/Swiss-Prot条目进行注释.
- PROSITE在SARS-CoV-2研究中发挥了重要作用,开发了病毒蛋白域的新工具和配置文件.
研究的目的:
- 更新和增强PROSITE数据库及其相关工具,以改善蛋白质注释和研究.
- 纳入新的数据源,如AlphaFold结构和本体,以更准确地定义域边界和特征注释.
- 提高PROSITE搜索功能的性能和可用性,特别是用于大规模的生物数据分析.
主要方法:
- 利用现有的PROSITE工具,为SARS-CoV-2蛋白域开发新的配置文件/ProRules.
- 将ChEBI本体学和Rhea参考词汇整合到ProRule中,用于化学联体和生物化学反应注释.
- 使用AlphaFold预测的3D结构来定义域边界并增强ScanProsite可视化.
- 重写和优化多核处理器的pfsearch代码,使用新的启发式来提高性能.
主要成果:
- PROSITE为SARS-CoV-2研究做出了贡献,并确定了POU2F和NF-κB转录因子协调器之间的联系.
- 现在ProRule包括了ChEBI和Rhea词汇,改善了对联体和反应的注释.
- AlphaFold结构用于域边界定义,ScanProsite允许对这些结构进行可视化.
- 对于现代硬件的速度和效率,pfsearch代码得到了显著的优化.
结论:
- PROSITE继续发展,为蛋白质域分析提供了必不可少的资源,并支持像病毒学这样的关键研究领域.
- 集成AlphaFold结构和增强的搜索功能提高了PROSITE注释的准确性和实用性.
- 持续的发展确保PROSITE仍然是生物信息学社区的一种强大而有效的工具.
更多相关视频
相关概念视频
Protein Families
16.6K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
16.6K
Protein Families
4.1K
4.1K
Protein-protein Interfaces
14.4K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
14.4K
Protein-Protein Interfaces
4.4K
4.4K
Conservation of Protein Domains Over Different Proteins
13.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
13.9K
Protein Networks
4.4K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.4K


