In silico prediction of chemical genotoxicity using machine learning methods and structural alerts
Defang Fan1, Hongbin Yang1, Fuxing Li1
1Shanghai Key Laboratory of New Drug Design , School of Pharmacy , East China University of Science and Technology , Shanghai 200237 , China . Email: gxliu@ecust.edu.cn ; Email: ytang234@ecust.edu.cn ; ; Tel: +86-21-64250811.
Abstract:
Genotoxicity tests can detect compounds that have an adverse effect on the process of heredity. The in vivo micronucleus assay, a genotoxicity test method, has been widely used to evaluate the presence and extent of chromosomal damage in human beings. Due to the high cost and laboriousness of experimental tests, computational approaches for predicting genotoxicity based on chemical structures and properties are recognized as an alternative. In this study, a dataset containing 641 diverse chemicals was collected and the molecules were represented by both fingerprints and molecular descriptors. Then classification models were constructed by six machine learning methods, including the support vector machine (SVM), naïve Bayes (NB), k-nearest neighbor (kNN), C4.5 decision tree (DT), random forest (RF) and artificial neural network (ANN). The performance of the models was estimated by five-fold cross-validation and an external validation set. The top ten models showed excellent performance for the external validation with accuracies ranging from 0.846 to 0.938, among which models Pubchem_SVM and MACCS_RF showed a more reliable predictive ability. The applicability domain was also defined to distinguish favorable predictions from unfavorable ones. Finally, ten structural fragments which can be used to assess the genotoxicity potential of a chemical were identified by using information gain and structural fragment frequency analysis. Our models might be helpful for the initial screening of potential genotoxic compounds.
Insights
Computational models can predict genotoxicity, identifying chemicals that damage heredity. This study developed accurate machine learning models to screen for potential genotoxic compounds, reducing the need for costly experimental tests.
Area of Science:
- Computational toxicology
- Cheminformatics
- Genetics
Background:
- Genotoxicity tests assess hereditary damage, but are costly and labor-intensive.
- Computational methods offer an alternative for predicting genotoxicity from chemical structures.
Purpose of the Study:
- To develop and validate machine learning models for predicting chemical genotoxicity.
- To identify structural fragments associated with genotoxic potential.
Main Methods:
- Collected a dataset of 641 chemicals, represented by fingerprints and molecular descriptors.
- Constructed classification models using six machine learning algorithms (SVM, NB, kNN, DT, RF, ANN).
- Evaluated model performance using five-fold cross-validation and an external validation set.
Main Results:
- Top models achieved high accuracy (0.846-0.938) on external validation.
- Pubchem_SVM and MACCS_RF models demonstrated reliable predictive ability.
- Identified ten structural fragments indicative of genotoxicity potential.
Conclusions:
- Machine learning models provide an effective approach for predicting chemical genotoxicity.
- These models can aid in the initial screening of potential genotoxic compounds.
- Identified structural fragments can assist in assessing genotoxicity risk.
More Related Videos
Related Concept Videos
Methods of Sterilization II: Chemical Methods
Using chemical sterilization rather than heat to clean out equipment is recommended. It eradicates and removes all bacteria,...
Predicting Molecular Geometry
Chemical Formulas
Machines
A free-body diagram of the...
Types of Chemical Bonds
Machines: Problem Solving II


