Related Experiment Video
Updated: Jan 17, 2026

Constructing and Visualizing Models using Mime-based Machine-learning Framework
Published on: July 22, 2025
Evaluating Machine Learning Models for Molecular Property Prediction: Performance and Robustness on
Hosein Fooladi1,2,3, Thi Ngoc Lan Vu1,2,3, Miriam Mathea4
1Department of Pharmaceutical Sciences, Division of Pharmaceutical Chemistry, Faculty of Life Sciences, University of Vienna, Josef-Holaubek-Platz 2, 1090 Vienna, Austria.
Machine learning models for molecular property prediction perform differently on out-of-distribution (OOD) data. Scaffold splitting shows good performance, while similarity clustering is challenging, impacting model selection for real-world applications.
Area of Science:
- * Cheminformatics
- * Computational Chemistry
- * Machine Learning
Background:
- * Machine learning models are widely used for predicting molecular properties.
- * Performance evaluation typically uses in-distribution (ID) data, but real-world use requires out-of-distribution (OOD) data.
- * Assessing model performance on OOD data is crucial for reliable predictions in novel chemical spaces.
Purpose of the Study:
- * To investigate and evaluate machine learning model performance on OOD molecular data.
- * To define OOD data generation strategies in molecular property prediction.
- * To analyze the relationship between in-distribution (ID) and out-of-distribution (OOD) performance.
Main Methods:
- * Evaluated 14 machine learning models, including random forests and graph neural networks (GNNs).
- * Utilized eight datasets and ten splitting strategies for OOD data generation.
- * Employed Bemis-Murcko scaffolds and UMAP-based clustering (ECFP4 fingerprints) for OOD splitting.
Main Results:
- * Bemis-Murcko scaffold splitting showed models performed well, similar to random splitting.
- * UMAP-based chemical similarity clustering presented the most challenging OOD scenario.
- * The correlation between ID and OOD performance varied significantly with splitting strategy (Pearson's r ~0.9 for scaffolds, ~0.4 for clusters).
Conclusions:
- * OOD data generation strategy critically influences model performance and ID-OOD correlation.
- * Scaffold-based splitting is less challenging than similarity-based clustering for OOD evaluation.
- * Model selection requires careful consideration of OOD performance aligned with specific application domains.
Related Concept Videos
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Predicting Reaction Outcomes
Predicting Molecular Geometry
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...
Mechanistic Models: Compartment Models in Algorithms for Numerical Problem Solving
In individual population analyses, different algorithms are employed, such as Cauchy's method, which uses a...

