Related Experiment Video
Updated: Aug 24, 2025

A Protocol for Computer-Based Protein Structure and Function Prediction
Published on: November 3, 2011
Assessment of the Generalization Abilities of Machine-Learning Scoring Functions for Structure-Based Virtual
Hui Zhu1,2, Jincai Yang2, Niu Huang1,2
1Tsinghua Institute of Multidisciplinary Biomedical Research, Tsinghua University, Beijing, China102206, China.
Machine-learning scoring functions (MLSFs) struggle with cross-target generalization in structure-based virtual screening (SBVS). Performance declines when tested on diverse targets, highlighting the need for robust evaluation methods like Pfam-clustering.
Area of Science:
- Computational chemistry
- Drug discovery
- Bioinformatics
Background:
- Structure-based virtual screening (SBVS) relies on accurate scoring functions to predict protein-ligand interactions.
- Machine-learning scoring functions (MLSFs) are increasingly used but their generalization across different protein targets remains a challenge.
Purpose of the Study:
- To develop and apply a standardized pocket Pfam-based clustering (Pfam-cluster) approach to assess the cross-target generalization ability of MLSFs.
- To evaluate the performance of 12 typical MLSFs using various cross-validation strategies.
Main Methods:
- A novel Pfam-cluster approach was introduced to evaluate MLSF generalization by focusing on local binding pocket domains.
- Twelve MLSFs were tested using random cross-validation (Random-CV), sequence similarity-based cross-validation (Seq-CV), and Pfam-based cross-validation (Pfam-CV).
- Interpretable analysis focused on features influencing predictions, particularly buried solvent-accessible surface area (SASA).
Main Results:
- All evaluated MLSFs demonstrated decreased performance when transitioning from Random-CV to Seq-CV and Pfam-CV, indicating limited generalization capacity.
- MLSF predictions on novel targets were significantly influenced by buried SASA-related features, favoring larger protein-ligand interfaces.
- The random forest (RF)-Score achieved good performance in Random-CV when incorporating both buried SASA features and target-specific patterns.
Conclusions:
- The Pfam-cluster approach is recommended for a more rigorous assessment of MLSF generalization ability.
- Caution is advised regarding the features learned by MLSFs, as they may not generalize well to unseen targets.
- Future MLSF development should consider incorporating features that capture both local pocket characteristics and broader target-specific patterns for improved cross-target performance.
More Related Videos
10:29Quantitative Structure-Activity Relationship, Activity Prediction, and Molecular Dynamics of Non-nucleotide Reverse Transcriptase Inhibitors
Published on: May 9, 2025
05:08Application of I TASSER, trRosetta, UCSF Chimera, HADDOCK server, and HEX loria for De Novo and In Silico Design of Proteins
Published on: July 8, 2025