接近:用于氨基酸关系的神经嵌入
Daniel Olson1, Thomas Colligan2, Daphne Demekas2
1Department of Computer Science, University of Montana, Missoula, MT 59812, United States.
Bioinformatics (Oxford, England)
|July 15, 2025
概括
氨基酸关系的神经嵌入 (NEAR) 提供比目前的方法更快,更准确的蛋白质同质检测. 这种新的方法提高了速度,并减少了蛋白质序列数据库搜索的虚假标签.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习在基因组学中的应用
背景情况:
- 蛋白质语言模型 (PLM) 是有前途的,但面临着速度和准确性的限制.
- 经典的序列对齐方法已经建立,但可以超越它们的性能.
- 高效的同质性搜索对于大规模的蛋白质注释至关重要.
研究的目的:
- 介绍氨基酸关系的神经嵌入 (NEAR) 以改进蛋白质同质性搜索.
- 提高识别同源蛋白序列的速度和准确性.
- 评估与最先进的方法相比NEAR的性能及其作为预先过器的实用性.
主要方法:
- 开发了NEAR,一种使用神经表示学习和ResNet嵌入模型的方法.
- 使用可信序列对齐引导的对比学习训练模型.
- 采用了残留级k-NN搜索和邻近聚合的管道来识别候选者.
主要成果:
- 与基准数据集上最先进的PLM相比,NEAR的准确性大大提高.
- NEAR 显示了更低的内存需求和更快的嵌入和搜索速度.
- NEAR作为高速预过器,性能优于HMMER3的预过器等现有方法.
结论:
- 在蛋白质同质检测方面,NEAR提供了显著的进步,平衡了速度和准确性.
- 该方法显示了具有增强灵敏度的独立同质检测的潜力.
- 对于敏感的注释管道,NEAR 作为一个有效和快速的预过器.
更多相关视频
07:08Optimization of Synthetic Proteins: Identification of Interpositional Dependencies Indicating Structurally and/or Functionally Linked Residues
Published on: July 14, 2015
7.4K
09:47Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
1.3K
相关概念视频
Amino acids
91.6K
Amino acids are the monomers that comprise proteins. Each amino acid has the same fundamental structure, which consists of a central carbon atom, or the alpha (α) carbon, bonded to an amino group (NH2), a carboxyl group (COOH), and to a hydrogen atom. Every amino acid also has another atom or group of atoms bonded to the central atom known as the R group. There are 20 common amino acids present in proteins, each with a different R group. Variation in the amino acid sequence is responsible...
91.6K
Protein Networks
4.1K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.1K
Evolutionary Relationships through Genome Comparisons
6.2K
Genome comparison is one of the excellent ways to interpret the evolutionary relationships between organisms. The basic principle of genome comparison is that if two species share a common feature, it is likely encoded by the DNA sequence conserved between both species. The advent of genome sequencing technologies in the late 20th century enabled scientists to understand the concept of conservation of domains between species and helped them to deduce evolutionary relationships across diverse...
6.2K
Structure of Amines
2.7K
The hybridized nitrogen atom in amines possesses a lone pair of electrons and is bound to three substituents with a bond angle of around 108°, which is less than the tetrahedral angle of 109.5°. However, the C–N–H bond angle is slightly larger at 112°, with a carbon–nitrogen bond length of 147 pm. This carbon–nitrogen bond length of of amines is longer than the carbon–oxygen bond of alcohols (143 pm) but shorter than alkanes’...
2.7K
tRNA Activation
20.0K
Aminoacyl-tRNA synthetases are present in both eukaryotes and bacteria. Though eukaryotes have 20 different aminoacyl-tRNA synthetases to couple to 20 amino acids, many bacteria do not have genes for all of these aminoacyl-tRNA synthetases. Despite this, they still use all 20 amino acids to synthesize their proteins. For instance, some bacteria do not have the gene encoding the enzyme that couples glutamine with its partner tRNA. In these organisms, one enzyme adds glutamic acid to all of the...
20.0K
Protein Organization
145.1K
Overview
145.1K
