对准对准对蛋白质细胞下定位预测的准确性的影响
Maryam Gillani1, Gianluca Pollastri1
1School of Computer Science, University College Dublin (UCD), Dublin, Ireland.
Proteins
|November 22, 2024
概括
多个序列对齐显著改善了深度学习模型,用于预测蛋白质细胞下定位. 将进化信息纳入生物信息学预测中可以提高模型的准确性和可靠性.
科学领域:
- 生物信息学是一种生物信息学.
- 计算生物学 计算生物学
- 机器学习 机器学习
背景情况:
- 序列对齐在生物信息学中至关重要,用于识别相似性和推断生物关系.
- 准确的蛋白质细胞下定位预测对于理解蛋白质功能和细胞过程至关重要.
研究的目的:
- 评估序列对齐对深层卷积神经网络对蛋白质亚细胞局部化预测性能的影响.
- 为了比较训练有或没有多重序列对齐 (MSAs) 的模型的准确性.
主要方法:
- 利用了不同深度和宽度的深度N-to-1卷积神经网络.
- 在没有对齐的模型中使用一次热编码来进行序列表示.
- 使用PSI-BLAST生成的MSA用于具有对齐的模型以捕获进化信息.
主要成果:
- 结合MSA的模型显示,与没有对齐的模型相比,MSA的平均性能改善约为15.82%.
- 使用对齐实现的最高精度大约比没有对齐提高了15.16%.
- 调整精度与不同类别的模型可靠性和预测一致性正相关.
结论:
- 序列对齐显著提高了蛋白质亚细胞定位的深度学习模型的准确性和可靠性.
- 这些发现表明,通过MSA结合进化信息是改善生物信息学预测工具的好处.
- 这项研究为生物信息学家提供了更准确的预测模型,可能减少了对广泛实验验证的需求.
相关概念视频
Conserved Binding Sites
4.2K
Many proteins’ biological role depends on their interactions with their ligands, small molecules that bind to specific locations on the protein known as ligand-binding sites. Ligand-binding sites are often conserved among homologous proteins as these sites are critical for protein function.
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
4.2K
Directing Proteins to the Rough Endoplasmic Reticulum
7.1K
The organelle-specific signaling sequences direct proteins synthesized in the cytosol to their final destination like ER, mitochondria, peroxisomes, etc. Some of the proteins directed to ER are then trafficked via vesicles to other organelles within the cell or the extracellular environment through the Golgi complex. For example, the rough ER synthesizes soluble proteins for transportation to the lysosomes or secretion out of the cell. It can also synthesize transmembrane proteins that can...
7.1K
Nuclear Localization Signals and Import
5.5K
Proteins targeted to the nucleus carry short stretches of amino acid sequences called the nuclear localization signal or NLS. Classical nuclear localization signals are of two types: monopartite and bipartite NLS. Monopartite classical NLS (cNLS) consists of a single cluster of 4-8 amino acids. Bipartite cNLS consists of two clusters of 2-3 amino acids and a 9-12 residue long proline-rich linker bridging the two clusters. Signal clusters are rich in positively charged amino acids such as...
5.5K
Tail-anchoring of Proteins in the ER Membrane
3.1K
Tail-anchored, or TA, proteins are estimated to make up to 3-5% of membrane proteins found in the eukaryotic cell. Such proteins have a single transmembrane domain located approximately 30 amino acid residues upstream from the C-terminal end. As a result, the signal recognition particle (SRP) cannot guide a TA protein to the ER membrane for cotranslational insertion. Hence, they are integrated into the ER membrane post-translationally using their C-terminal end as the anchor. TA proteins...
3.1K
Protein-protein Interfaces
12.5K
Many proteins form complexes to carry out their functions, making protein-protein interactions (PPIs) essential for an organism's survival. Most PPIs are stabilized by numerous weak noncovalent chemical forces. The physical shape of the interfaces determines the way two proteins interact. Many globular proteins have closely-matching shapes on their surfaces, which form a large number of weak bonds. Additionally, many PPIs occur between two helices or between a surface cleft and a...
12.5K
Protein Organization
6.2K
Proteins are polymers of amino acid residues. They are versatile and responsible for different cellular functions, including DNA replication, molecular transport, catalysis, and structural support. Proteins have a hierarchical structure comprising at least three levels of organization: primary, secondary, and tertiary structure. Some large proteins have a quaternary structure where individual protein subunits are linked together.
The primary structure of a protein is its amino acid sequence....
The primary structure of a protein is its amino acid sequence....
6.2K


