A Comparison of Machine Learning Approaches for predicting Hepatotoxicity potential using Chemical Structure and
Tia Tate1, Grace Patlewicz1, Imran Shah1
1Center for Computational Toxicology and Exposure, Office of Research and Development, U.S. Environmental Protection Agency, Research Triangle Park, North Carolina 27709, USA.
Computational Toxicology (Amsterdam, Netherlands)
|July 12, 2024
Summary
Addressing bias in animal toxicity data is crucial for accurate machine learning (ML) predictions. This study explored sampling methods to balance toxicity datasets, finding that tailored ML workflows are essential for reliable toxicity predictions.
Area of Science:
- Toxicology
- Computational Chemistry
- Bioinformatics
Background:
- Animal toxicity testing is resource-intensive, creating a bottleneck for substance assessment.
- Existing toxicity datasets for training machine learning (ML) models are often biased due to the selection of substances likely to cause toxicity.
- This bias can significantly impact the predictive performance of ML models used in toxicity assessment.
Purpose of the Study:
- To investigate the impact of data balancing strategies on the predictive performance of ML models for hepatotoxicity.
- To evaluate various sampling approaches for mitigating class imbalance in toxicity datasets.
- To compare the performance of different ML algorithms and feature sets in predicting toxicity outcomes.
Main Methods:
- Utilized supervised ML workflows with chemical structure and/or transcriptomic data to predict hepatotoxicity.
- Applied various sampling approaches (over-sampling, under-sampling) to balance imbalanced in vivo toxicity data.
- Evaluated 18 study-toxicity outcome combinations across multiple ML models, including Artificial Neural Networks, Random Forests, and k-Nearest Neighbour (k-NN).
Main Results:
- Unbalanced data for chronic liver effects yielded a mean CV F1 score of 0.735.
- Over-sampling approaches decreased performance to 0.639, while under-sampling resulted in 0.523.
- Developmental liver toxicity showed improved performance with over-sampling (mean CV F1 of 0.234) compared to unbalanced data (0.089).
- Model performance was influenced by dataset characteristics, model type, and balancing strategy.
Conclusions:
- Class imbalance significantly affects ML model performance in toxicity prediction.
- Tailoring ML workflows, considering data balancing and starting with simpler classifiers, is recommended for accurate toxicity prediction.
- The choice of balancing approach and ML model is critical for optimizing predictive accuracy across different toxicity endpoints.


