Related Experiment Video
Updated: May 1, 2026

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Structure-Aware Compound-Protein Affinity Prediction via Graph Neural Networks with Group Lasso Regularization
Zanyu Shi1, Yang Wang2, Pathum M Weerawarna3
1Department of Biostatistics & Health Data Science, Indiana University Fairbanks School of Public Health, Indianapolis, 46202 IN, USA.
Abstract:
Explainable artificial intelligence approaches accelerate drug discovery by improving molecular representation learning, identifying key molecular structures, and rationalizing drug property prediction. However, developing end-to-end explainable models for structure-activity relationship modeling in target-specific compound property prediction remains challenging due to the limited availability of compound-protein interaction data for individual targets and the fact that small changes in chemical substituents or local structural motifs can lead to large differences in molecular properties. Thus, optimally leveraging structural and property information and identifying key moieties related to compound-protein affinity for specific targets is essential. We propose a framework implementing graph neural networks (GNNs) to leverage property and structure information from pairs of molecules with activity cliffs targeting specific proteins to predict compound-protein affinity (i.e., half-maximal inhibitory concentration, IC50) and explain property differences. To enhance model explainability, we trained GNNs with structure-aware loss functions using group lasso and sparse group lasso regularizations, which prune and highlight molecular subgraphs relevant to activity differences. We applied this framework to the activity cliff data of molecules targeting 6 tyrosine-protein kinases across Src, Abl, and Tec families, as well as anaplastic lymphoma kinase. Integrating common- and uncommon-node information with sparse group lasso improves molecular property prediction for specific protein targets, as evidenced by lower root mean square errors and higher Pearson's correlation coefficients. Applying regularizations also enhances feature attribution for GNNs by boosting graph-level global direction scores and improving atom-level coloring accuracy. These advances strengthen model interpretability in drug discovery pipelines, particularly in identifying critical molecular substructures in lead optimization.
Related Concept Videos
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Protein Networks
Protein-protein Interfaces
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Ligand Binding Sites

