使用协同方法进行高级蛋白质分类和功能预测的全面框架:集成双光谱分析,机器学习和深度学习
Hiam Alquran1, Amjed Al Fahoum1, Ala'a Zyout1
1Hijjawi Faculty for Engineering Technology, Biomedical Systems and Informatics Engineering Department, Yarmouk University, Irbid, Jordan.
PloS one
|December 14, 2023
概括
使用双光谱分析和深度学习的新方法准确地分类蛋白质家族,优于传统技术,更好地了解蛋白质功能和疾病. 这推动了蛋白质生物学和药物发现.
科学领域:
- 计算生物学 计算生物学
- 生物信息学是一种生物信息学.
- 机器学习在生物学中的应用
背景情况:
- 蛋白质是参与疾病过程的必不可少的细胞组成部分.
- 准确的蛋白质家族分类对于进化研究,功能预测和治疗点识别至关重要.
- 像序列和结构对齐等传统方法往往不足以有效识别蛋白质家族.
研究的目的:
- 开发一种更有效,更准确的蛋白质特征提取和分类方法.
- 克服现有的技术在识别蛋白质家族的局限性.
- 增强用于蛋白质家族识别的分类指标.
主要方法:
- 一种新的方法,集成双频谱特征,深度学习 (卷积神经网络) 和机器学习算法.
- 蛋白质序列以数值表示,并使用双光谱分析进行分析.
- 深度学习模型提取特征,然后进行强大的特征选择以进行分类.
主要成果:
- 拟议的方法在蛋白质家族识别方面显著优于传统方法.
- 增强的分类指标证明了双频谱和深度学习策略的优越有效性.
- 该方法的有效性在众多蛋白质数据集中得到验证.
结论:
- 新的双光谱和深度学习方法为蛋白质家族识别提供了更高的精度和效率.
- 这些进展在科学学科中具有广泛的适用性,有助于理解蛋白质功能和疾病.
- 这些发现支持制药创新,并加深我们对蛋白质在健康和疾病中的作用的了解.
相关概念视频
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Protein Families
15.4K
Protein families are groups of homologous proteins; that is, they have similarities in amino acid sequences and three-dimensional structures. Protein families usually occur because of gene duplication, where an additional copy of a gene is inserted into the genome of an organism. Mutations that change the amino acids but still allow the protein to be properly synthesized, will lead to new protein family members. If these new proteins contain similar amino acids in key...
15.4K
Genome Annotation and Assembly
18.9K
The genome refers to all of the genetic material in an organism. It can range from a few million base pairs in microbial cells to several billion base pairs in many eukaryotic organisms. Genome assembly refers to the process of taking the DNA sequencing data and putting it all back together in a correct order to create a close representation of the original genome. This is followed by the identification of functional elements on the newly assembled genome, a process called genome annotation.
18.9K


