Related Experiment Video
Updated: Jan 23, 2026

07:15
Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
7.4K
Error Tolerance of Machine Learning Algorithms across Contemporary Biological Targets
Thomas M Kaiser1, Pieter B Burger2,3
1St Peter's College, University of Oxford, New Inn Hall St, Oxford OX1 2DL, UK. thomas.kaiser@spc.ox.ac.uk.
Molecules (Basel, Switzerland)
|June 7, 2019
Summary
Machine learning models for drug discovery can tolerate significant data inaccuracies. Computational methods like Free Energy Perturbation (FEP+) can generate useful datasets when experimental data is scarce.
Area of Science:
- Computational chemistry
- Machine learning
- Drug discovery
Background:
- Machine learning (ML) models are crucial for predicting drug properties.
- ML model efficacy relies on accurate and abundant data, which are often scarce.
- Investigating data accuracy limitations can reveal alternative data sources for ML.
Purpose of the Study:
- To assess the impact of data inaccuracies on ML model performance in drug development.
- To explore the utility of datasets with known error profiles, such as those from Free Energy Perturbation (FEP+).
- To determine if ML models can be trained on less accurate data for drug property prediction.
Main Methods:
- Introduced varying proportions of error into high-accuracy datasets across diverse targets (kinases, GPCR, polymerases, proteases).
- Evaluated the retrospective accuracy decay of Naïve Bayes Network, Random Forest, and Probabilistic Neural Network models.
- Tested the performance of ML models trained on data simulating FEP+ error profiles.
Main Results:
- Naïve Bayes Networks tolerated up to 39% training set error before losing predictivity.
- Random Forests tolerated an average of 29% training set error.
- Probabilistic Neural Networks showed lower tolerance, losing predictivity at around 20% error.
- Naïve Bayes Networks and Random Forests successfully utilized FEP+-like error profiles.
Conclusions:
- ML models demonstrate notable tolerance to data inaccuracies in drug property prediction.
- Computational methods with known error distributions, like FEP+, can serve as valuable data sources.
- This approach may reduce reliance on extensive and costly in vitro experimental datasets for ML model generation.
Keywords:
FEPNaïve Bayes NetworkNeural NetworkRandom Forestanaplastic lymphoma kinase (ALK)cheminformaticsdrug discoveryerrormachine learningMore Related Videos
Related Concept Videos
Trial and Error and Algorithm
403
A problem-solving strategy is a plan of action used to find a solution. Different strategies have distinct action plans. Trial and error involves trying different solutions until one works. For instance, to fix a broken printer, you might check ink levels, ensure the paper tray isn't jammed, and verify the printer's connection to your laptop. This method can be time-consuming but is commonly used. Thomas Edison, for example, used trial and error to find a suitable filament for the light...
403
Contemporary Psychology
2.2K
Psychology explores human behavior and mental processes through various lenses, each offering unique insights. This overview examines key subfields, including biopsychology, evolutionary, developmental, personality, and social psychology, highlighting their approaches and contributions to understanding complex human behaviors.
Biopsychology
Biopsychology, also known as biological psychology or behavioral neuroscience, focuses on the biological underpinnings of behavior and mental processes. It...
Biopsychology
Biopsychology, also known as biological psychology or behavioral neuroscience, focuses on the biological underpinnings of behavior and mental processes. It...
2.2K
Systematic Error: Methodological and Sampling Errors
9.9K
In the case of systematic errors, the sources can be identified, and the errors can be subsequently minimized by addressing these sources. According to the source, systematic errors can be divided into sampling, instrumental, methodological, and personal errors.
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
Sampling errors originate from improper sampling methods or the wrong sample population. These errors can be minimized by refining the sampling strategy. Defective instruments or faulty calibrations are the sources of instrumental...
9.9K
Fundamental Attribution Error
13.7K
According to some social psychologists, people tend to overemphasize internal factors as explanations—or attributions—for the behavior of other people. They tend to assume that the behavior of another person is a trait of that person, and to underestimate the power of the situation on the behavior of others. They tend to fail to recognize when the behavior of another is due to situational variables, and thus to the person’s state. This erroneous assumption is...
13.7K
Random Error
9.1K
Random or indeterminate errors originate from various uncontrollable variables, such as variations in environmental conditions, instrument imperfections, or the inherent variability of the phenomena being measured. Usually, these errors cannot be predicted, estimated, or characterized because their direction and magnitude often vary in magnitude and direction even during consecutive measurements. As a result, they are difficult to eliminate. However, the aggregate effect of these errors can be...
9.1K
Margin of Error
7.1K
The margin of error is also called the maximum error of an estimate. The margin of error is the maximum possible or expected difference between the observed sample parameter value and the actual population parameter value. For proportion, it is the maximum difference between the value of sample proportion obtained from the data and the true value of population proportion. As the true value of the population parameter is not known, the margin of error is calculated using the sample statistic.
7.1K

