Related Experiment Videos
Machine Learning-Driven Prediction of Dose-Linear Pharmacokinetics: Utilizing Molecular Descriptors to Guide
Elizabeth S Levy1, Karen E Samy2, Andrea Anelli3
1Department of Synthetic Molecule Pharmaceutical Sciences, Genentech Inc, South San Francisco, California 94080, United States.
Abstract:
Early stage preclinical drug formulation development can be impacted by limited materials and time, especially when making rapid, data-driven decisions is critical. An important parameter to consider is dose linearity, as this can guide the selection between determining if a conventional or enabled formulation strategy is necessary to achieve the expected target exposure as the dose is escalated. This work explores the application of machine learning models to predict dose linearity for compounds administered to the mouse species based on the molecular properties. We utilize a diverse set of models, including Random Forest, XGBoost, LASSO, Neural Networks, and TabPFN, using two distinct data sets for training: a full data set including measured and calculated descriptors and a more accessible data set containing only the RDKit descriptors. Cross-validation was applied with random and scaffold splits, and model performance was evaluated based on a test data set containing 10% of the full data set that was held out during training. Our analysis indicated that the ML models trained with the full descriptors were comparable to the SMILES-only derived descriptors with RDKit. Additionally, while most models achieved comparable performance, LASSO was prioritized due to the interpretability, feature selection, and resistance to overfitting smaller data sets. With a random split and RDKit-only descriptors, the LASSO model resulted in 90.9% precision in the test set and 88.0% accuracy. Precision was optimized to minimize false positives, as incorrectly determining that a compound would not be less-than-dose proportional, where a conventional vehicle would be sufficient, has a higher cost than false negatives. To further evaluate the model's capacity to apply to novel chemical compounds, a scaffold split was tested. A gap in performance was seen with a precision of 89.3% in training and 80.9% in testing, highlighting the challenges in predicting outcomes with new, chemically unique architectures. A significant finding of this work is the robust performance of the models trained solely on accessible RDKit descriptors, demonstrating that a tool can be developed without more resource-intensive measured data. Overall, this research focuses on a data-driven approach based on dose linearity prediction for whether a conventional, nonsolubilizing vehicle is sufficient or if an enabled formulation would be required to achieve the target exposures. The application to forecast a compound's dose linearity risk with only the chemical structures provides a valuable tool to guide early formulation strategy and accelerate compound evaluation and preclinical development during the drug discovery phase.
Related Concept Videos
Determination of Multiple Dosing Parameters: Loading and Maintenance Doses
Nonlinear Pharmacokinetics: Overview
Nonlinearity can arise due to the saturation of plasma protein-binding or...
Pharmacokinetic–Pharmacodynamic Relationship: Problems
Pharmacokinetic–Pharmacodynamic Relationship: Model Components
Dosage Regimens: Partial Pharmacokinetic Parameters
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.