Related Experiment Video
Updated: Jun 18, 2025

12:58
Characterizing Individual Protein Aggregates by Infrared Nanospectroscopy and Atomic Force Microscopy
Published on: September 12, 2019
9.8K
Massive experimental quantification of amyloid nucleation allows interpretable deep learning of protein aggregation
Mike Thompson1, Mariano Martín2, Trinidad Sanmartín Olmo2
1Systems and Synthetic Biology, Centre for Genomic Regulation, The Barcelona Institute for Science and Technology (BIST), Barcelona, Spain.
Biorxiv : the Preprint Server for Biology
|July 29, 2024
Summary
Researchers experimentally quantified amyloid nucleation for over 100,000 protein sequences. This led to CANYA, a new neural network model that accurately predicts protein aggregation from sequence.
Area of Science:
- Biochemistry
- Computational Biology
- Genetics
Background:
- Protein aggregation is implicated in over fifty human diseases.
- Existing methods for predicting protein aggregation are limited by small, biased datasets.
- Accurate prediction of protein aggregation is crucial for both understanding disease and biotechnology applications.
Purpose of the Study:
- To address the data shortage in protein aggregation prediction.
- To develop an accurate computational model for predicting amyloid nucleation from protein sequence.
- To provide an interpretable deep learning model for analyzing protein aggregation propensities.
Main Methods:
- Experimentally quantified amyloid nucleation propensity for over 100,000 random protein sequences.
- Trained a novel convolution-attention hybrid neural network (CANYA) on the generated dataset.
- Utilized genomic neural network interpretability techniques to analyze the model's decision-making process.
Main Results:
- The large-scale dataset revealed limitations in existing protein aggregation prediction methods.
- The developed CANYA model accurately predicts amyloid nucleation from protein sequence.
- Interpretability analysis provided insights into the learned sequence grammar governing amyloid nucleation.
Conclusions:
- Massive experimental analysis of random sequence spaces is a powerful approach for biological discovery.
- CANYA offers an interpretable and robust computational tool for predicting protein amyloid nucleation.
- The findings advance our ability to understand and engineer proteins, with implications for disease and biotechnology.

