Related Experiment Video
Updated: Oct 4, 2026

Human Pluripotent Stem Cell Based Developmental Toxicity Assays for Chemical Safety Screening and Systems Biology Data Generation
Published on: June 17, 2015
Artificial Intelligence and machine learning in toxicogenomics for drug safety assessment: a critical review
Baseerat Fatima1, Minal Zaka2, Maham Ashraf1
1Department of Pharmacy, The University of Faisalabad, Faisalabad, Pakistan.
Introduction:
Predicting drug-induced toxicity remains one of the principal bottlenecks in drug development: in vivo and in vitro assays are resource-intensive, and their concordance with human-specific toxicity is often limited. Toxicogenomics addresses this gap by linking chemical exposure to gene-expression and pathway-level changes, offering mechanistic insight that conventional endpoints cannot provide.
Objectives:
Artificial intelligence (AI) and machine learning (ML) increasingly analyze these high-dimensional datasets drawn from resources such as Tox21, ToxCast, TG-GATEs, LINCS, Drug Matrix, CTD and DrugBank to model complex, non-linear relationships between molecular signatures and organ-specific toxicities such as hepatotoxicity, nephrotoxicity and cardiotoxicity. Although previous surveys have catalogued AI/ML applications in predictive toxicology, few have critically examined how dataset composition, validation strategy, interpretability and regulatory requirements jointly determine whether these predictions can be trusted. This review addresses that gap.
Results:
Its central finding is that AI/ML has substantially expanded the analytical capacity of toxicogenomics, but reported predictive performance commonly AUC values spanning roughly 0.70 to above 0.90 for the same clinical endpoint does not by itself establish biological validity, external generalizability, reproducibility, or regulatory acceptance. This variability is caused by chemical biases, species biases in public datasets, as well as low sample size and imbalanced samples, nearuniversal use of internal crossvalidation, while independent external testing is preferred, and batch effects that do not always get corrected. Explainable AI methods (SHAP, LIME, attention mechanisms) may be employed to find potential biomarkers and more generally to improve the mechanistical transparency (but are not regulatory-grade mechanistical evidence, as they are not proven by independent experiments and are sometimes sensitive to minor changes in the input).
Conclusion:
We feel that methodological discipline, including the use of standardized benchmark datasets, the external validation of the datasets as standard practice, instead of optional, balanced reporting of performance, the characterization of applicability domains, and proactive experimental confirmation of computationally-derived biomarkers, are important areas requiring future improvements. Without improving data quality, benchmarking, and reporting standards, the use of multi-omics approaches, further validated explainable architectures, and careful, careful reporting of provenance with generative AI may provide plausible paths for increasing the translational and regulatory potential of AI-driven toxicogenomics in the future.
Related Concept Videos
Toxicokinetics: Overview
Drug Toxicity: Overview
Drug toxicity: Idiosyncratic Reactions
Drug Toxicity: Dose-Dependent Reactions
Toxicity Testing in Animals
Pharmacogenomics: Identification of New Drug Targets
