Related Experiment Video
Updated: Jan 25, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
A decision-theoretic approach to the evaluation of machine learning algorithms in computational drug discovery
Oliver P Watson1, Isidro Cortes-Ciriano1,2, Aimee R Taylor3,4
1Goring on Thames, Evariste Technologies Ltd., RG8 9AL UK.
Motivation:
Artificial intelligence, trained via machine learning (e.g. neural nets, random forests) or computational statistical algorithms (e.g. support vector machines, ridge regression), holds much promise for the improvement of small-molecule drug discovery. However, small-molecule structure-activity data are high dimensional with low signal-to-noise ratios and proper validation of predictive methods is difficult. It is poorly understood which, if any, of the currently available machine learning algorithms will best predict new candidate drugs.
Results:
The quantile-activity bootstrap is proposed as a new model validation framework using quantile splits on the activity distribution function to construct training and testing sets. In addition, we propose two novel rank-based loss functions which penalize only the out-of-sample predicted ranks of high-activity molecules. The combination of these methods was used to assess the performance of neural nets, random forests, support vector machines (regression) and ridge regression applied to 25 diverse high-quality structure-activity datasets publicly available on ChEMBL. Model validation based on random partitioning of available data favours models that overfit and 'memorize' the training set, namely random forests and deep neural nets. Partitioning based on quantiles of the activity distribution correctly penalizes extrapolation of models onto structurally different molecules outside of the training data. Simpler, traditional statistical methods such as ridge regression can outperform state-of-the-art machine learning methods in this setting. In addition, our new rank-based loss functions give considerably different results from mean squared error highlighting the necessity to define model optimality with respect to the decision task at hand.
Availability And Implementation:
All software and data are available as Jupyter notebooks found at https://github.com/owatson/QuantileBootstrap.
Supplementary Information:
Supplementary data are available at Bioinformatics online.
More Related Videos
12:18A Machine Learning Approach to Design an Efficient Selective Screening of Mild Cognitive Impairment
Published on: January 11, 2020
08:49Incorporating Target Protein Structure Flexibility and Dynamics in Computational Drug Discovery Using Ensemble-Based Docking Analysis
Published on: June 20, 2025
Related Concept Videos
Theoretical Approaches to Psychological Disorder
Biological approach
The biological approach posits that internal, organic factors are the primary causes of such disorders. This perspective emphasizes brain structure and function, genetic predispositions, and neurotransmitter imbalances. For example, schizophrenia has been associated with both genetic...
Drug Discovery: Overview
Theoretical Foundations of Nursing Practice
Theories provide a perspective to assess patients' conditions and organize data and methods. They also assist in analyzing and interpreting information. They represent a...
Machines
A free-body diagram of the...
Trial and Error and Algorithm
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...