Related Experiment Video
Updated: Jan 27, 2026

Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
Classical scoring functions for docking are unable to exploit large volumes of structural and interaction data.
Hongjian Li1,2, Jiangjun Peng3, Pavel Sidorov4
1SDIVF R&D Centre, Hong Kong Science Park, Sha Tin, New Territories, Hong Kong.
Machine learning scoring functions (SFs) improve accuracy with more diverse training data, unlike classical SFs. Random forest and XGBoost models learn effectively from dissimilar protein-ligand complexes, enhancing predictive power.
Area of Science:
- Computational chemistry
- Drug discovery
- Bioinformatics
Background:
- Scoring functions (SFs) are crucial for predicting protein-ligand binding affinity.
- Random Forest (RF)-based SFs show improved accuracy with more training data, unlike classical SFs.
- The impact of training-test similarity on SF performance is not well understood.
Purpose of the Study:
- To systematically investigate how training-test similarity affects classical and machine learning SF accuracy.
- To determine if machine learning SFs can generalize from dissimilar training data.
- To develop and evaluate a novel SF using Extreme Gradient Boosting (XGBoost).
Main Methods:
- Evaluated classical (X-Score) and machine learning (RF-Score-v3) SFs.
- Assessed performance using three similarity metrics: protein structure, protein sequence, and ligand structure.
- Developed and tested a new XGBoost-based SF (XGB-Score).
Main Results:
- Classical SFs did not improve accuracy with increased training set size or similarity.
- RF-Score-v3 outperformed X-Score, especially when trained on dissimilar complexes.
- XGB-Score demonstrated improved accuracy with training set size and outperformed other SFs.
Conclusions:
- Machine learning SFs, particularly RF and XGBoost, benefit significantly from diverse training data, including dissimilar complexes.
- The ability to learn from dissimilar data is key to the superior performance of modern SFs.
- XGBoost presents a promising new approach for developing accurate SFs in drug discovery.
More Related Videos
Related Concept Videos
Structural Protein Function
Collagen, the most abundant protein in mammals, is found throughout the body. In connective tissue, such as skin, ligaments, and tendons, it provides tensile strength and elasticity. In bones and teeth, it mineralizes to...
Structural Protein Function
Fruit Development, Structure, and Function
Structure and Function of Erythrocytes
The erythrocyte plasma membrane is associated with proteins such as spectrin, which forms a flexible cytoplasmic meshwork. This meshwork allows erythrocytes to twist, turn, become cup-shaped, and regain their biconcave shape as they pass through narrow capillaries. Additionally, erythrocytes can form...
Structure and Function of Platelets
Platelets are continually replenished, circulating in the bloodstream for 9-12 days before being removed by phagocytes, primarily in the spleen. A microliter of circulating blood contains between 150,000 and 450,000...
Classical Conditioning
Ivan Pavlov observed that dogs...

