Related Experiment Video
Updated: Dec 30, 2025

09:47
Author Spotlight: Advancing Alzheimer's Research – Exploring Early Detection and Multi-Omics Approaches
Published on: December 15, 2023
1.6K
Validating the validation: reanalyzing a large-scale comparison of deep learning and machine learning models for
Matthew C Robinson1, Robert C Glen2,3, Alpha A Lee4
1Department of Physics, J J Thomson Avenue, Cambridge, CB3 0HE, UK.
Journal of Computer-Aided Molecular Design
|January 22, 2020
Summary
Support vector machines show competitive performance against deep learning for drug discovery bioactivity prediction. Researchers suggest using precision-recall curves alongside ROC curves for better virtual screening evaluation.
Area of Science:
- Computational chemistry
- Bioinformatics
- Machine learning in drug discovery
Background:
- Machine learning accelerates drug discovery, but consistent benchmarking is challenging.
- Numerous new machine learning methods require robust validation strategies.
Purpose of the Study:
- To re-evaluate machine learning model performance for bioactivity prediction.
- To question the suitability of standard metrics in virtual screening.
- To propose improved evaluation approaches for drug discovery models.
Main Methods:
- Reanalysis of a large-scale machine learning model comparison dataset.
- Numerical experiments to assess performance metrics.
- Evaluation of scaffold-split nested cross-validation for uncertainty estimation.
Main Results:
- Support vector machines demonstrate performance competitive with deep learning methods.
- Area under the receiver operating characteristic curve may have limited relevance for virtual screening.
- Area under the precision-recall curve is recommended alongside ROC curves.
Conclusions:
- SVMs remain a viable option in machine learning for drug discovery.
- Precision-recall curves offer valuable insights for virtual screening.
- Careful consideration of validation strategies is crucial for reliable model assessment.

