Related Experiment Video
Updated: Jul 4, 2026

High Content Screening Analysis to Evaluate the Toxicological Effects of Harmful and Potentially Harmful Constituents (HPHC)
Published on: May 10, 2016
A Comprehensive Study to Compare Different Compound Representations for Predicting Carcinogenicity In Vivo.
Iuri Barbosa Pereira1, Rogerio Salvini2, Eloisa Dutra Caldas1
1Laboratório de Toxicologia, Faculdade de Ciências da Saúde, Universidade de Brasília, Brasília, Federal District, Brazil.
Predicting chemical carcinogenicity computationally is crucial for risk assessment. Combining molecular descriptors with structural alerts offers the most reliable machine learning models for predicting in vivo carcinogenicity.
Area of Science:
- Computational toxicology
- Chemical risk assessment
- Machine learning in drug discovery
Background:
- In vivo carcinogenicity testing is resource-intensive and ethically complex.
- Machine learning (ML) models offer a promising alternative for carcinogenicity prediction.
- The impact of different molecular representations on ML model performance requires systematic investigation.
Purpose of the Study:
- To evaluate the influence of molecular embeddings, classical descriptors, and structural alerts on ML models for predicting in vivo carcinogenicity.
- To compare the predictive performance of various molecular representation strategies.
- To identify optimal approaches for regulatory-relevant carcinogenicity modeling.
Main Methods:
- Assembled a dataset of 2090 compounds with in vivo rodent carcinogenicity data from five toxicological databases.
- Represented compounds using classical molecular descriptors, structural alerts, SMILES-derived embeddings, and hybrid combinations.
- Benchmarked 24 ML classifiers using 10-fold cross-validation, assessing performance via accuracy, precision, recall, F1-score, and AUC-ROC.
Main Results:
- Hybrid representations combining molecular descriptors and structural alerts consistently yielded the best predictive performance across ML models.
- Molecular embeddings served as valuable complementary features but did not outperform classical representations.
- Chemically interpretable, expert-driven descriptors, especially those with genotoxic alerts, are vital for robust carcinogenicity modeling.
Conclusions:
- Combining molecular descriptors with structural alerts is a superior strategy for building accurate ML models for carcinogenicity prediction.
- Classical molecular representations remain essential, with embeddings offering supplementary predictive power.
- These findings support the use of interpretable descriptors in regulatory toxicology for chemical safety assessment.
More Related Videos
Related Concept Videos
Mutagenicity and Carcinogenicity
Toxicity Testing in Animals
Mouse Models of Cancer Study
The development of transgenic, knockout, and knock-in mice has led to an exponential increase in their use as model organisms in research,...
Types of Biopharmaceutical Studies: Controlled and Non-Controlled Approaches
Non-controlled studies, commonly employed for initial exploration, lack a control group, rendering them susceptible to biases and external influences. In contrast, controlled...

