Related Experiment Video
Updated: Aug 5, 2026

A Bilingual Computational Workflow for Identifying Potential PLK1 Inhibitors in American Sign Language and English
Published on: April 3, 2026
An ML-Based QSAR Web Server for KEAP1 Inhibitor Bioactivity Prediction: Composite-Score-Driven Training and Advanced
1Department of Chemistry and Biochemistry, Louise Dilworth Davis College of Science & Engineering, Texas Christian University, Fort Worth, Texas 76129, United States.
Abstract:
Cardiovascular disorders and neurodegenerative diseases are among the leading causes of death worldwide, with oxidative stress being a prominent etiological factor in their development. Lowering chronic oxidative stress has been proposed as a strategy for improving or treating these conditions. The body does this naturally through the KEAP1:NRF2 pathway in healthy individuals. Inhibiting the KEAP1 regulatory protein releases the NRF2 transcription factor, leading to the biosynthesis of the antioxidant proteins. Novel molecules that activate this pathway have recently been proposed as potential mechanisms to halt diseases driven by oxidative stress, but this approach has not yet reached clinical translation. To advance this approach to drug development, there is a strong need to rapidly identify KEAP1-specific molecules. Incorporating machine-learning tools into the drug development process reduces the risk of failure. Hence, this study presents a quantitative structure-activity relationship-based machine-learning model that can predict the potential of novel KEAP1 inhibitors before synthesis and biological evaluation. To achieve this goal, molecular fingerprints of KEAP1 inhibitors retrieved from ChEMBL and BindingDB were generated by using PaDEL, Mordred, and RDKit. Subsequently, these fingerprints were screened using a novel composite-score-based feature selection method, and the resulting features were then used to train 30 models. Their performances were rigorously evaluated and ranked using the coefficient of determination (R 2), root-mean-square error (RMSE), and mean absolute error (MAE). The best model was further validated using the concordance correlation coefficient (CCC), external validation (Q F1 2, Q F2 2, and Q F3 2), k-fold testing, and y-scrambling. On completion of the study, CatBoost demonstrated the highest predictive power with strong R 2 values for the test (0.8373) and training (0.9548) sets. A significant improvement was observed in CCC, external validation, cross-validation, and the y-scrambling test. This model was finally deployed as a Web server, which is freely available to researchers.