Related Experiment Video
Updated: Sep 11, 2026

In Silico Modeling Method for Computational Aquatic Toxicology of Endocrine Disruptors: A Software-Based Approach Using QSAR Toolbox
Published on: August 28, 2019
From QSAR to deep learning: an interpretable comprehensive pipeline with a read-across approach for mutagenicity
Dimitra-Danai Varsou1,2, Andreas Tsoumanis2,3, Eleonora Marta Longhin4
1NovaMechanics MIKE Piraeus 18545 Greece varsou@novamechanics.com afantitis@novamechanics.com.
Abstract:
Assessing the mutagenicity of chemical compounds is essential for ensuring their safe handling and use, thereby minimizing potential health risks. New approach methodologies (NAMs), including computational approaches, provide non-animal alternatives for testing novel materials and chemicals. This study highlights the potential of in silico NAMs, which can contribute to the development of novel Safe and Sustainable by Design (SSbD) chemicals and substances by identifying potentially hazardous ones at an early stage. Emphasis is given to the mutagenicity prediction based on data from the Ames (bacterial gene mutation) test curating them to consider stereo-specific input whenever necessary. A consensus strategy integrating different chemical representations (molecular fingerprints, 2D and 3D descriptors and molecular graphs) and modelling methods, i.e., Quantitative Structure-Activity Relationship (QSAR), read-across and deep learning models, is employed to predict the mutagenic profile of chemical compounds. In this course, a XGBoost model is developed based on molecular descriptors and a graph convolutional neural networks model to classify compounds as mutagens and non-mutagens based on the Ames test data. The devised read-across methodology is based on a guided-k-Nearest Neighbours scheme (guided-kNN) where two different molecular representations (molecular fingerprints and descriptors) are considered for neighbour selection and predictions generation. The mutagenicity predictions from the three models are integrated in a majority voting scheme to enhance the overall predictive accuracy (83% in external validation) and reduce individual model biases. Interpretation of the descriptors involved in prediction is performed through explainable AI (XAI) methods to provide insight to the mutagenicity mechanism. To enhance the interpretability of the XAI-derived insights and reinforce user confidence in the models' predictions, the involved descriptors are mapped to key events leading to mutations within the Adverse Outcome Pathway (AOP) networks. Apart from the development of reliable and interpretable mutagenicity models, emphasis is given on delivering a pipeline for the generation of 3D descriptors that can be used as the basis for future cheminformatics models. To support transparency and reproducibility of the results of our work, the curated mutagenicity dataset used for modelling is disseminated through the ChemPharos database (https://db.chempharos.eu/datasets/Datasets.zul?datasetID=ds18), the modelling steps are documented following the standardized Modelling Data (MODA) guidelines and the consensus model is freely available via the Enalos Cloud platform (https://www.enaloscloud.novamechanics.com/insight/polis/), to facilitate virtual screening of novel compounds.
More Related Videos
16:02Demonstration of the Sequence Alignment to Predict Across Species Susceptibility Tool for Rapid Assessment of Protein Conservation
Published on: February 10, 2023
14:34A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026