下一篇: 神经嵌入用于氨基酸关系的神经嵌入
bioRxiv : the preprint server for biology
|February 3, 2025
概括
NEAR是一种用于蛋白质同质性搜索的新方法,比现有的蛋白质语言模型 (PLM) 提供了更好的速度和准确性. 这种神经网络方法增强了蛋白质数据库的搜索,使其更快,更可靠地识别相关的蛋白质序列.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 结构生物信息学 结构生物信息学
背景情况:
- 蛋白质语言模型 (PLM) 在替代蛋白质数据库搜索的传统序列对齐方法方面表现有前途.
- 目前的PLM通常比基于对齐的工具慢,产生更多的假阳性.
- 有效和准确的同质检测对于理解蛋白质功能和进化至关重要.
研究的目的:
- 推出NEAR,一种基于神经网络的方法,旨在提高在大数据库中搜索同类蛋白质的速度和准确性.
- 评估与最先进的PLM和现有的预先过方法对比NEAR的性能.
主要方法:
- NEAR使用了一个ResNet嵌入模型,该模型通过可信的序列对齐指导的对比学习进行训练.
- 它计算了蛋白质序列的每残余嵌入.
- 同类候选物被识别使用一个管道,涉及残留级 k-最近邻居 (k-NN) 搜索和邻居聚合.
主要成果:
- 与最先进的PLM相比,NEAR在远程同类和诱的基准上显示了大幅提高的准确性.
- 该方法显示了更低的内存需求和更快的嵌入和搜索速度.
- 作为配置文件隐藏马尔科夫模型 (pHMM) 搜索的预过器,NEAR至少比HMMER3的预过器快5倍,并且性能优于钉子工具中的预过器.
结论:
- NEAR在蛋白质同质性搜索方面取得了重大进展,提高了速度和准确性.
- 它作为高速预过器的有效性提高了敏感的注释管道,优于当前的方法.
- NEAR为独立同质检测提供了一个有价值的替代方案,其灵敏度可能比标准对齐方法更高.
更多相关视频
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
1.3K
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
942
相关概念视频
Amino acids
87.7K
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible...
87.7K
Protein Networks
3.9K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
3.9K
Evolutionary Relationships through Genome Comparisons
5.7K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
5.7K
Structure of Amines
2.4K
The hybridized nitrogen atom in amines possesses a lone pair of electrons and is bound to three substituents with a bond angle of around 108°, which is less than the tetrahedral angle of 109.5°. However, the C–N–H bond angle is slightly larger at 112°, with a carbon–nitrogen bond length of 147 pm. This carbon–nitrogen bond length of of amines is longer than the carbon–oxygen bond of alcohols (143 pm) but shorter than alkanes’...
2.4K
tRNA Activation
18.9K
Aminoacyl-tRNA synthetases are present in both eukaryotes and bacteria. Though eukaryotes have 20 different aminoacyl-tRNA synthetases to couple to 20 amino acids, many bacteria do not have genes for all of these aminoacyl-tRNA synthetases. Despite this, they still use all 20 amino acids to synthesize their proteins. For instance, some bacteria do not have the gene encoding the enzyme that couples glutamine with its partner tRNA. In these organisms, one enzyme adds glutamic acid to all of the...
18.9K
Protein Organization
136.4K
Overview
136.4K
