Related Experiment Videos
Revisiting ADMET prediction reliability under real-world challenges in the foundation model era
Donghai Zhao1,2, Yuchen Zhu1, Zhenxing Wu1
1Department of Clinical Pharmacy, College of Pharmaceutical Sciences, The First Affiliated Hospital, Zhejiang University School of Medicine, Zhejiang University, Hangzhou, 310058, Zhejiang, China.
Foundation models transform AI drug discovery. This study benchmarks models on data scarcity, generalization, and activity cliffs, finding tabular models excel in OOD scenarios and ensembles improve robustness for better druggability prediction.
Area of Science:
- Artificial Intelligence in Drug Discovery
- Machine Learning for Cheminformatics
- Computational Drug Design
Background:
- AI-driven drug discovery is shifting towards foundation models.
- Existing molecular foundation models lack systematic benchmarking for real-world challenges.
- Druggability prediction needs robust evaluation across diverse scenarios.
Purpose of the Study:
- Establish a benchmark for evaluating AI models in drug discovery.
- Assess limitations and frontiers of molecular and tabular foundation models.
- Provide guidance for selecting and improving druggability prediction models.
Main Methods:
- Developed a benchmark with four challenges: data scarcity/OOD generalization, class imbalance, beyond Rule of 5 (bRo5) generalization, and activity cliffs.
- Performed large-scale benchmarking of molecular foundation models (KPGT, Uni-Mol) and tabular foundation models (TabPFNv2).
- Evaluated ensemble strategies and the molecular property landscape roughness index.
Main Results:
- Tabular foundation models show superior generalization in few-shot and OOD settings.
- Ensemble methods with undersampling effectively handle extreme class imbalance.
- Graph Neural Networks (GNNs) like KPGT perform well in bRo5 chemical spaces with sufficient data.
- All models struggle with predicting activity cliffs, indicating a persistent challenge.
Conclusions:
- Ensembling diverse, high-performance models enhances predictive robustness.
- The molecular property landscape roughness index is a useful tool for model selection.
- Findings offer practical guidance for evaluating and selecting AI models for druggability prediction.
Related Concept Videos
Accuracy, limits, and approximation
Accuracy is defined as the closeness of the measured value to the true or actual value. In engineering mechanics, repeated measurements are taken during theoretical or experimental analyses to ensure that the result is precise and accurate.
The accuracy of any solution is based on the...
Improving Translational Accuracy
Improving Translational Accuracy
Reliability and Validity
Olefin Metathesis Polymerization: Acyclic Diene Metathesis (ADMET)
Similar to cross-metathesis, ADMET also involves the formation of metallacyclobutane intermediate by [2+2] cycloaddition of one of the double bonds of a terminal diene with...
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
The...