蛋白质中氨基酸模式的可解释机器学习:一种统计整体方法
Anna Braghetto1,2, Enzo Orlandini1,2, Marco Baiesi1,2
1Department of Physics and Astronomy, University of Padova, Via Marzolo 8, 35131 Padua, Italy.
Journal of chemical theory and computation
|August 8, 2023
概括
无监督的机器学习模型揭示了意想不到的蛋白质结构特性. 对受限制的博尔茨曼机器进行集体分析,揭示了阿尔法螺旋和β片中的氨基酸作用.
科学领域:
- 计算生物学是一种计算生物学.
- 生物信息学是一种生物信息学.
- 在蛋白质科学中的机器学习
背景情况:
- 了解蛋白质的二次结构 (α螺旋和β片) 对于破译蛋白质的功能至关重要.
- 无监督机器学习为揭示复杂生物数据中隐藏的模式提供了强大的工具.
- 可解释模型对于从机器学习分析中获得有意义的生物学见解至关重要.
研究的目的:
- 开发和应用机器学习模型的整体分析,以便更好地解释蛋白质序列数据.
- 研究蛋白质二次结构边界的氨基酸序列中的信息含量和模式.
- 通过机器学习识别氨基酸的新特性及其对蛋白质二次结构形成的贡献.
主要方法:
- 使用受限制的博尔兹曼机器 (RBM) 作为机器学习组合的核心组件.
- 应用集体分析来巩固来自多个机器学习模型的解释.
- 分析了RBM的学习重量,以确定显著的氨基酸模式和特性.
主要成果:
- 有限制的博尔茨曼机器有效地压缩来自螺旋/片末端的五个氨基酸序列的信息.
- 确定了特定的氨基酸倾向及其在阿尔法螺旋的两模式中的作用.
- 发现His和Thr对阿尔法螺旋的两性影响很小,而存在阿拉-丰富的螺旋.
- 突出了proline在启动螺旋的位置和作用,经常取代极性/充电残留物.
- 在Glu,Asp,Val,Leu,Ile和Phe中观察到强烈的两性标志物,与有效的疏水性有关.
结论:
- 整体机器学习为解释复杂的生物序列数据提供了一个强大的框架.
- 这项研究揭示了特定氨基酸在蛋白质二次结构形成中的细微,以前未知的作用.
- 这些发现有助于通过可解释的人工智能更深入地了解蛋白质折叠和结构功能关系.
相关概念视频
Proteomics
7.4K
A proteome is the entire set of proteins that a cell type produces. We can study proteomes using the knowledge of genomes because genes code for mRNAs, and the mRNAs encode proteins. Although mRNA analysis is a step in the right direction, not all mRNAs are translated into proteins.
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
Proteomics is the study of proteomes' function. It involves the large-scale systematic study of the proteome to denote the protein complement expressed by a genome. Scientist Mark Wilkins coined the term...
7.4K
Peptide Identification Using Tandem Mass Spectrometry
6.5K
Tandem mass spectrometry, also known as MS/MS or MS2, is an analytical technique that employs two mass analyzers. Essentially it is a series of mass spectrometers that helps isolate a particular biomolecule and then helps study its chemical properties.
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
This technique helps gather information regarding the protein from which the peptide was obtained and to study the peptides’ amino acid sequence. Identifying peptides from a complex mixture is an important component of the growing field of...
6.5K
Conservation of Protein Domains Over Different Proteins
10.9K
Protein domains are small structurally independent units that are part of a single amino acid chain. Although these domains are often structurally independent, they may rely on synergistic effects to perform their functions as part of a larger protein. Protein domains may be conserved within the same organism, as well as across different organisms.
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
10.9K
Protein Networks
4.0K
An organism can have thousands of different proteins, and these proteins must cooperate to ensure the health of an organism. Proteins bind to other proteins and form complexes to carry out their functions. Many proteins interact with multiple other proteins creating a complex network of protein interactions.
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
4.0K


