Related Experiment Video
Updated: Jul 16, 2025

Kinase Inhibitor Screening In Self-assembled Human Protein Microarrays
Published on: October 23, 2019
Poor Generalization by Current Deep Learning Models for Predicting Binding Affinities of Kinase Inhibitors
Wern Juin Gabriel Ong1,2, Palani Kirubakaran1, John Karanicolas1
1Cancer Signaling & Microenvironment Program, Fox Chase Cancer Center, Philadelphia, PA 19111.
Convolutional Neural Networks (CNNs) show promise for predicting drug-receptor binding affinities. However, current models exhibit significant data leakage, failing to generalize to new data and offering no advantage over simpler methods.
Area of Science:
- Computational chemistry
- Drug discovery
- Machine learning in pharmacology
Background:
- Neural networks are increasingly used to predict drug-receptor binding affinities, potentially accelerating drug discovery.
- Existing models using protein kinase sequences and SMILES strings report accurate quantitative inhibition predictions.
- Concerns remain regarding the generalizability of these neural network models to novel, unseen data.
Approach:
- A Convolutional Neural Network (CNN) model was developed, similar to previously reported architectures.
- The CNN was evaluated on four standard datasets for inhibitor/kinase binding predictions.
- Model performance was assessed under random data splitting versus inhibitor-grouped data splitting.
Key Points:
- The CNN achieved comparable performance to prior models when data points were randomly split between training and testing sets.
- Performance drastically decreased when all data for a specific inhibitor was grouped in either the training or testing set, indicating information leakage.
- The model demonstrated no significant generalization ability, performing no better than simple baseline models like tokenized SMILES or nearest neighbor predictions.
Conclusions:
- The observed lack of generalization suggests current CNN models do not truly learn molecular interactions for binding affinity prediction.
- Information leakage, particularly when data is not properly segregated, inflates performance metrics.
- Richer, structure-based molecular encodings are necessary for developing models capable of reliable prospective predictions of new drug candidates.
Related Concept Videos
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Protein-protein Interfaces
Physiological Pharmacokinetic Models: Assumption with Protein Binding
The Equilibrium Binding Constant and Binding Strength
Protein-Drug Binding: Mechanism and Kinetics
Various forces drive these interactions, including hydrogen bonds, hydrophobic interactions, ionic bonds, electrostatic interactions, and van der Waals forces. These bonds enable drugs to bind to specific sites on proteins,...

