Related Experiment Video
Updated: Jun 11, 2025

Evaluation of the Impact of Protein Aggregation on Cellular Oxidative Stress in Yeast
Published on: June 23, 2018
Enhancing protein aggregation prediction: a unified analysis leveraging graph convolutional networks and active
Jiwon Sun1, JunHo Song1, Juo Kim1
1School of Mechanical Engineering, Soongsil University 369 Sangdo-ro, Dongjak-gu Seoul 06978 Republic of Korea kmin.min@ssu.ac.kr.
This study developed a Graph Convolutional Network (GCN) model to predict protein aggregation (PA) propensity, achieving high accuracy. An active learning approach further enhanced efficiency in identifying proteins prone to aggregation.
Area of Science:
- Biochemistry and Molecular Biology
- Computational Biology
- Structural Biology
Background:
- Protein aggregation (PA) is implicated in neurodegenerative diseases like Alzheimer's and Parkinson's.
- Understanding PA requires insights into aggregation-prone regions (APRs) and structural interactions.
- Computational methods, especially machine learning, offer efficient alternatives to experimental PA studies.
Purpose of the Study:
- To develop a Graph Convolutional Network (GCN) model for accurate protein aggregation (PA) score prediction.
- To leverage expanded datasets from Protein Data Bank (PDB) and AlphaFold2.0 for improved model training.
- To evaluate an active learning strategy for efficient identification of proteins with high PA propensity.
Main Methods:
- Constructed a GCN model utilizing an enhanced dataset derived from PDB and AlphaFold2.0.
- Calculated PA propensity using AGGRESCAN3D 2.0 and refined PDB data by separating multi-polypeptide chains.
- Incorporated 22,774 Homo sapiens sequences from AlphaFold2.0 after sequence similarity comparison.
Main Results:
- The trained GCN model achieved a high coefficient of determination (R 2) of 0.9849 and a low mean absolute error (MAE) of 0.0381 for PA prediction.
- The active learning approach demonstrated superior performance with an MAE of 0.0291 in expected improvement.
- Active learning identified 99% of target proteins by exploring only 29% of the search space.
Conclusions:
- The developed GCN model shows significant promise for predicting protein aggregation susceptibility.
- The active learning strategy enhances the efficiency of identifying proteins prone to aggregation.
- This work advances computational tools for PA prediction, with potential applications in disease diagnosis and therapy.
More Related Videos
07:35A Knowledge Graph Approach to Elucidate the Role of Organellar Pathways in Disease via Biomedical Reports
Published on: October 13, 2023
06:50Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein-protein Interfaces
Amyloid Fibrils
Amyloid deposits were observed as early as 1639 in the liver and the spleen. In 1854, Rudolph Virchow performed iodine staining,...
Protein Organization
The primary structure of a protein is its amino acid sequence....