Related Experiment Video
Updated: Jan 9, 2026

Quantitative Structure-Activity Relationship, Activity Prediction, and Molecular Dynamics of Non-nucleotide Reverse Transcriptase Inhibitors
Published on: May 9, 2025
Machine Learning-Enhanced Quantitative Structure-Activity Relationship Modeling for DNA Polymerase Inhibitor
Samuel Kakraba1, Srinivas Ayyadevara2, Aayire Yadem Clement3
1Department of Biostatistics and Data Science, Celia Scott Weatherhead School of Public Health and Tropical Medicine, Tulane University, 1440 Canal St, New Orleans, LA, 70112, United States, 1 5049882475.
Background:
Cisplatin resistance remains a significant obstacle in cancer therapy, frequently driven by translesion DNA synthesis mechanisms that use specialized polymerases such as human DNA polymerase η (hpol η). Although small-molecule inhibitors such as PNR-7-02 have demonstrated potential in disrupting hpol η activity, current compounds often lack sufficient potency and specificity to effectively combat chemoresistance. The vastness of chemical space further limits traditional drug discovery approaches, underscoring the need for advanced computational strategies such as machine learning (ML)-enhanced quantitative structure-activity relationship (QSAR) modeling.
Objective:
This study aimed to develop and validate ML-augmented QSAR models to accurately predict hpol η inhibition by indole thio-barbituric acid analogs, with the goal of accelerating the discovery of potent and selective inhibitors that could overcome cisplatin resistance.
Methods:
A curated library of 85 indole thio-barbituric acid analogs with validated hpol η inhibition data was used, excluding outliers to ensure data integrity. Molecular descriptors spanning 1D to 4D were computed in MAESTRO, resulting in 220 features. In total, 17 ML algorithms, including random forest, extreme gradient boosting (XGBoost), and neural networks, were trained using 80% of the data for training and evaluated with 14 performance metrics. Robustness was ensured through hyperparameter optimization and 5-fold cross-validation.
Results:
Ensemble methods outperformed other algorithms, with random forest achieving near-perfect predictive performance (training mean square error=0.0002; R²=0.9999 and testing mean square error=0.0003; R²=0.9998). Shapley additive explanations analysis revealed that electronic properties, lipophilicity, and topological atomic distances were the most important predictors of hpol η inhibition. Linear models exhibited higher error rates, highlighting the nonlinear relationship between molecular descriptors and inhibitory activity.
Conclusions:
Integrating ML with QSAR modeling provides a robust framework for optimizing hpol η inhibition, offering both high predictive accuracy and biochemical interpretability. This approach accelerates the identification of potent selective inhibitors and represents a promising strategy for overcoming cisplatin resistance, thereby advancing precision oncology.
More Related Videos
22:10Multi-target Parallel Processing Approach for Gene-to-structure Determination of the Influenza Polymerase PB2 Subunit
Published on: June 28, 2013
07:38DNA Polymerase Activity Assay Using Near-infrared Fluorescent Labeled DNA Visualized by Acrylamide Gel Electrophoresis
Published on: October 6, 2017
Related Concept Videos
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Proofreading
Errors During Replication are Corrected by the DNA Polymerase...