Related Experiment Video
Updated: May 10, 2026

10:29
Quantitative Structure-Activity Relationship, Activity Prediction, and Molecular Dynamics of Non-nucleotide Reverse Transcriptase Inhibitors
Published on: May 9, 2025
Comparison of confirmed inactive and randomly selected compounds as negative training examples in support vector
Kathrin Heikamp1, Jürgen Bajorath
1LIMES Program Unit, Chemical Biology and Medicinal Chemistry, Department of Life Science Informatics, Rheinische Friedrich-Wilhelms-Universität, Dahlmannstr. 2, D-53113 Bonn, Germany.
Journal of Chemical Information and Modeling
|June 27, 2013
Summary
Choosing negative training data significantly impacts machine learning models for drug discovery. This study reveals that typical benchmarks overestimate support vector machine (SVM) performance in virtual screening compared to real-world applications.
Area of Science:
- Chemoinformatics
- Machine Learning
- Computational Chemistry
Background:
- The selection of negative training data is a critical yet understudied aspect of machine learning in chemoinformatics.
- Support Vector Machine (SVM) models are widely used for virtual screening in drug discovery.
- The composition of training datasets can significantly affect model performance and predictive accuracy.
Purpose of the Study:
- To investigate the influence of different negative training data sets and background databases on SVM model performance.
- To evaluate the impact of these choices on virtual screening outcomes and hit recall.
- To compare benchmark settings with more practical application scenarios for SVM-based virtual screening.
Main Methods:
- Derived target-directed SVM models using training sets with confirmed inactive molecules or randomly selected compounds as negative instances.
- Applied these SVM models to search diverse background databases, including biological screening data and random compound collections.
- Systematically analyzed the effect of negative data composition and background database choice on virtual screening results.
Main Results:
- Negative training data composition was found to systematically influence compound recall in virtual screening.
- The choice of background databases significantly impacted the search results and identified hits.
- Typical benchmark settings tend to overestimate the performance of SVM-based virtual screening.
Conclusions:
- The selection of negative training data and background databases are crucial factors that must be carefully considered in chemoinformatics.
- Virtual screening performance estimations based on standard benchmarks may not accurately reflect real-world applicability.
- Optimizing negative data selection and background database choice is essential for reliable and practical virtual screening in drug discovery.
