Related Experiment Video
Updated: Jun 30, 2025

Investigating Protein Sequence-structure-dynamics Relationships with Bio3D-web
Published on: July 16, 2017
Classifying protein kinase conformations with machine learning
Ivan Reveguk1, Thomas Simonson1
1Laboratoire de Biologie Structurale de la Cellule (CNRS UMR7654), Ecole Polytechnique, Palaiseau, France.
Machine learning models accurately classify protein kinase conformations, distinguishing active/inactive states and DFG motif positions. This aids in understanding kinase signaling and drug development by analyzing structural data.
Area of Science:
- Structural biology
- Computational biology
- Pharmacology
Background:
- Protein kinases are crucial in cell signaling and are significant drug targets.
- Kinase activity is regulated by conformational changes, particularly in the catalytic domain's activation loop and DFG motif.
- Accurate structural annotation of kinases is vital for targeted drug design but requires scalable and interpretable methods.
Purpose of the Study:
- To develop and validate interpretable machine learning models for automated annotation of protein kinase structures.
- To classify kinase structures based on their active/inactive states and DFG motif conformations (DFG-in, DFG-out, other).
- To identify key structural features driving kinase conformational states and assess the accuracy of predicted structures from tools like AlphaFold2.
Main Methods:
- Collected and curated a diverse dataset of protein kinase catalytic domain sequences and structures.
- Clustered structures based on DFG conformation and manually annotated them for training.
- Developed ensemble decision tree models, initially using 1692 structural variables, to classify kinase states and DFG conformations.
Main Results:
- The active/inactive classification model achieved 99.9% accuracy on 3289 structures.
- The DFG conformation model achieved >99.8% accuracy on 8826 structures.
- Identified key structural variables, primarily near the activation loop, that are critical for classification, providing insights into conformational preferences.
Conclusions:
- Interpretable machine learning models offer a robust and automated approach for large-scale structural annotation of protein kinases.
- These models accurately classify kinase conformations, aiding in the understanding of signaling pathways and drug targeting.
- Analysis of AlphaFold2-predicted structures revealed discrepancies in DFG-in proportions compared to the Protein Data Bank, highlighting the utility of these models for evaluating predicted structures.
Related Concept Videos
Protein Kinases and Phosphatases
Protein kinases
Many proteins in the cell are regulated by phosphorylation, the addition of a phosphate group. A family of enzymes called kinases...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Protein Organization
The primary structure of a protein is its amino acid sequence....
Conservation of Protein Domains Over Different Proteins
A limited set of protein domains often duplicate and recombine during evolution. These domains can be organized in different combinations to...
MAPK Signaling Cascades
Protein-protein Interfaces

