从它们的进化概况来预测抗蛋白的存在
Nishant Kumar1, Shubham Choudhury1, Nisha Bajiya1
1Department of Computational Biology, Indraprastha Institute of Information Technology, New Delhi, India.
Proteomics
|September 21, 2024
概括
对抗蛋白 (AFP) 的准确预测对于医疗保健至关重要. 使用进化信息的新模型显著优于现有方法,在经过审查的蛋白质数据上实现了高精度.
科学领域:
- 生物化学和生物信息学
- 计算生物学 计算生物学
- 蛋白质科学 蛋白质科学
背景情况:
- 抗蛋白 (AFP) 在医学和生物技术中具有重要应用.
- 现有的AFP预测工具经常使用未经审查的蛋白质数据集,限制了它们的可靠性.
- 需要强大且经过验证的方法来预测AFP.
研究的目的:
- 评估现有的,并提出用于抗蛋白 (AFP) 预测的新计算方法.
- 使用验证的数据集开发一个高度准确和可靠的AFP预测模型.
- 为AFP预测提供一个用户友好的工具.
主要方法:
- 机器学习模型是使用基于组成的蛋白质特征来构建的.
- 在独立的,经过专家审查的数据集上评估了80个AFP和73个非AFP的表现,这些数据来自UniProt.
- 纳入了进化信息以提高模型性能.
- 研究了将机器学习与BLAST和基于动机的方法相结合的混合模型.
主要成果:
- 最初的机器学习模型实现了0.90的AUROC和0.69的MCC.
- 整合进化信息改善了模型性能,将AUROC提高到0.93.
- 最好的机器学习模型,利用进化信息,在独立的数据集上超过了所有现有方法.
- 混合模型没有超过最好的机器学习模型的性能.
结论:
- 结合进化信息的新型计算方法为防蛋白预测提供了更高的准确性.
- 开发的模型,AFPropred,为AFP识别提供了一个可靠和用户友好的解决方案.
- 这项研究为使用经过验证和专家审查的数据进行AFP预测建立了新的基准.
相关概念视频
Protein Families
15.3K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.3K
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Conservation of Protein Domains Over Different Proteins
10.8K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.8K
Multi-species Conserved Sequences
3.9K
Next-generation sequencing technologies have created large genomic databases of a variety of animals and plants. Ever since the human genome project was completed, scientists studied the genome of primates, mammals, and other phylogenetically distant living beings. Such large-scale studies have provided new insights into the evolutionary relationship between organisms.
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
Although the genome of each species varies greatly from each other, a few sequences are highly conserved. Such conserved...
3.9K
Gene Evolution - Fast or Slow?
7.0K
The genomes of eukaryotes are punctuated by long stretches of sequence which do not code for proteins or RNAs. Although some of these regions do contain crucial regulatory sequences, the vast majority of this DNA serves no known function. Typically, these regions of the genome are the ones in which the fastest change, in evolutionary terms, is observed, because there is typically little to no selection pressure acting on these regions to preserve their sequences.
In contrast, regions which code...
In contrast, regions which code...
7.0K


