Related Experiment Video
Updated: Jan 10, 2026

Author Spotlight: A Computational Approach to Decipher Amino Acid Preferences in Multispecific Protein-Protein Interactions
Published on: January 26, 2024
Benchmarking Sequence-Based Compound-Protein Interaction Prediction through Constructing a Debiased Data Set CDPN
Yang Hao1,2,3,4, Bo Li2,3, Daiyun Huang2,5
1Department of Hepatobiliary Surgery, Haikou Affiliated Hospital of Central South University Xiangya School of Medicine, Haikou 570208, P.R. China.
Abstract:
Accurate prediction of compound-protein interactions (CPIs) is critical for drug discovery, but existing data sets often suffer from biases that hinder model generalization. Here, we first highlighted that over-represented molecular scaffolds and imbalanced label distributions can lead to machine learning shortcuts. While existing debiasing approaches often compromise data set diversity, we present Clustering-based Down-sampling and Putative Negatives (CDPN), a novel protocol for constructing a debiased CPI benchmark. CDPN mitigates biases through compound Cluster-level Down-sampling and generates Putative Negatives from unexplored chemical spaces, ensuring balanced label distributions. Using CDPN, we systematically benchmark deep learning-based CPI models, with a particular focus on protein language models. Although systematic evaluation on PDBbind reveals critical limitations in attention interpretability, thorough ablation studies on the CDPN data set identify superior models such as KPGT-Ankh, which exhibits enhanced generalization and virtual screening performance. The top-performing models from benchmark were also integrated into DeepSEQreen, a no-code web server designed to facilitate community feedback and broader accessibility.
Related Concept Videos
Protein-protein Interfaces
Protein Networks
These interactions can be represented through maps depicting protein-protein interaction networks, represented as nodes and edges. Nodes are circles that are representative of a protein,...
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...

