Related Experiment Video
Updated: Apr 19, 2026

Nano-Differential Scanning Fluorimetry for Screening in Fragment-based Lead Discovery
Published on: May 16, 2021
Benchmarking methods and data sets for ligand enrichment assessment in virtual screening
Jie Xia1, Ermias Lemma Tilahun2, Terry-Elinor Reid2
1State Key Laboratory of Natural and Biomimetic Drugs, School of Pharmaceutical Sciences, Peking University, Beijing 100191, PR China; Molecular Modeling and Drug Discovery Core for District of Columbia Developmental Center for AIDS Research (DC D-CFAR), Laboratory of Cheminformatics and Drug Design, Department of Pharmaceutical Sciences, College of Pharmacy, Howard University, Washington, DC 20059, USA.
Virtual screening (VS) benchmarking sets can be biased. We developed a new algorithm to create maximum-unbiased benchmarking sets for ligand-based and structure-based VS, improving real-world assessment accuracy.
Area of Science:
- Computational chemistry
- Drug discovery
- Bioinformatics
Background:
- Retrospective virtual screening (VS) using benchmarking datasets is common for evaluating VS approaches.
- Existing benchmarking sets may not accurately reflect real-world screening libraries, leading to biased assessments.
- Key biases include analogue bias, artificial enrichment, and false negatives.
Purpose of the Study:
- To critically review the history and limitations of current VS benchmarking methods and datasets.
- To introduce a novel algorithm for constructing maximum-unbiased benchmarking sets for VS.
- To apply and validate this algorithm for structure-based and ligand-based VS targeting HDAC1, HDAC6, and HDAC8.
Main Methods:
- Comprehensive review of VS benchmarking history, datasets, and identified biases.
- Development of a new algorithm to generate maximum-unbiased benchmarking sets.
- Application of the algorithm to human histone deacetylase (HDAC) isoforms (HDAC1, HDAC6, HDAC8) and validation using leave-one-out cross-validation (LOO CV).
Main Results:
- Identified and categorized three primary biases in conventional VS benchmarking sets: analogue bias, artificial enrichment, and false negatives.
- Demonstrated the effectiveness of the developed algorithm in creating maximum-unbiased benchmarking sets.
- LOO CV confirmed the unbiased nature of the generated sets through property matching, ROC curves, and AUC values for HDAC targets.
Conclusions:
- Conventional VS benchmarking sets can introduce significant biases, potentially misrepresenting the true performance of VS methods.
- The newly developed algorithm effectively generates maximum-unbiased benchmarking sets suitable for both ligand-based and structure-based VS.
- This approach enhances the reliability of VS performance evaluation, particularly for drug discovery targets like HDACs.
More Related Videos
10:29Quantitative Structure-Activity Relationship, Activity Prediction, and Molecular Dynamics of Non-nucleotide Reverse Transcriptase Inhibitors
Published on: May 9, 2025
14:34A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
Related Concept Videos
The Equilibrium Binding Constant and Binding Strength
Ligand Binding Sites
Protein-ligand interactions are quite specific; even though numerous potential ligands surround a cellular protein at any given time, only a particular ligand can bind to that protein. Moreover, a ligand binds only to a dedicated area on the surface of the protein, known as the...
Ligand Binding Sites
Conserved Binding Sites
Binding sites are often located in large pockets, and if their location on a protein’s surface is unknown, it can be predicted using various approaches. The energetic method computationally...