Related Experiment Video
Updated: Jan 11, 2026

In Silico Modeling Method for Computational Aquatic Toxicology of Endocrine Disruptors: A Software-Based Approach Using QSAR Toolbox
Published on: August 28, 2019
Comparative Study of Molecular Descriptors and AI-Based Embeddings for Toxicity Prediction
1Division of Bioinformatics and Biostatistics, National Center for Toxicological Research, U.S. FDA, 3900 NCTR Rd, Jefferson, Arkansas 72079, United States.
AI language models show promise in predictive toxicology, outperforming traditional methods on specific datasets like ClinTox and DILIst. Molecular descriptors remain strong for multi-endpoint predictions, suggesting combined approaches for enhanced drug safety evaluation.
Area of Science:
- Computational chemistry
- Toxicology
- Artificial intelligence in drug discovery
Background:
- Accurate toxicity prediction is crucial for pharmaceutical development and regulatory safety.
- Traditional methods rely on molecular descriptor-based models.
- Emerging AI language models offer new approaches for chemical data analysis.
Purpose of the Study:
- To compare the performance of descriptor-based features versus AI language model embeddings for toxicity prediction.
- To evaluate models across Tox21, ClinTox, and DILIst datasets.
- To assess the utility of different data inputs (SMILES, chemical names, descriptions) for AI models.
Main Methods:
- Utilized descriptor-based features from Mordred and RDKit.
- Applied ten AI language models to generate embeddings from SMILES strings, chemical names, and descriptions.
- Employed logistic regression classifiers for toxicity prediction tasks.
- Evaluated performance using ROC-AUC on Tox21, ClinTox, and DILIst datasets.
Main Results:
- Mordred (descriptor-based) achieved the highest average ROC-AUC (0.855) on the Tox21 dataset.
- AI language models outperformed descriptor models on ClinTox and DILIst datasets.
- GPT-3 demonstrated superior performance on ClinTox (ROC-AUC 0.996) and DILIst (ROC-AUC 0.806) using textual data.
- MolBERT showed competitive performance on Tox21 (ROC-AUC 0.801) using SMILES embeddings.
Conclusions:
- AI language models show significant promise for predictive toxicology, especially for focused endpoints.
- Descriptor-based models remain effective for multi-endpoint predictions.
- Integrating molecular descriptors with textual embeddings could enhance predictive accuracy and model adaptability.
- This study highlights the potential of AI in revolutionizing drug safety assessment.
More Related Videos
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
09:01A High-throughput Assay for the Prediction of Chemical Toxicity by Automated Phenotypic Profiling of Caenorhabditis elegans
Published on: March 14, 2019