混合DBRpred:改进了基于DNA结合氨基酸的序列预测,使用结构复合体和无序蛋白质的注释
Jian Zhang1, Sushmita Basu2, Lukasz Kurgan2
1School of Computer and Information Technology, Xinyang Normal University, Xinyang 464000, PR China.
Nucleic acids research
|December 4, 2023
概括
目前的DNA结合残留预测剂在内在无序的蛋白质上表现不佳. 一个新的元模型,混合DBRpred,结合了顶级预测因素,以准确地识别结构化和无序蛋白质的DNA结合残留物.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 蛋白质科学 蛋白质科学
背景情况:
- 预测DNA结合残留物 (DBRs) 对于理解蛋白质功能至关重要.
- 现有的DBR预测器在结构化或内在无序的蛋白质数据上进行训练,导致性能差距.
研究的目的:
- 在不同蛋白质类型中实证分析当前DBR预测器的性能.
- 开发一种改进的元模型,用于在结构化和无序蛋白质中准确预测DBR.
主要方法:
- 对十个现有的DBR预测器进行实证性能分析.
- 基于深度变压器网络的元模型 (混合DBRpred) 的开发.
- 对混合DBRpred与现有工具和基线元预测器的验证.
主要成果:
- 结构训练的预测者在结构化蛋白质上表现出色,但在无序蛋白质上失败,反之亦然.
- 没有一个现有的预测器可以准确地识别两种蛋白质类型中的DBR.
- 与个人预测器和基线元模型相比,hybridDBRpred表现出卓越的准确性和减少的交叉预测.
结论:
- 当前的DBR预测器在应用于各种蛋白质结构时存在局限性.
- 混合DBRpred提供了一个强大的解决方案,用于在结构化和内在无序蛋白中准确识别DBR.
- 混合DBRpred网络服务器和源代码可供公众研究使用.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Protein Organization
6.5K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.5K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K


