Related Experiment Video
Updated: May 12, 2025

Quantitative Structure-Activity Relationship, Activity Prediction, and Molecular Dynamics of Non-nucleotide Reverse Transcriptase Inhibitors
Published on: May 9, 2025
Improved QSAR methods for predicting drug properties utilizing topological indices and machine learning models
Muhammad Shoaib Sardar1, Muhammad Shahid Iqbal2, Muhammad Mudassar Hassan3
1College of Mathematical Sciences, Harbin Engineering University, Harbin, People's Republic of China.
Abstract:
This research investigates the anticipated physicochemical and topological properties of compounds such as drug complexity (C), molecular weight (MW), and topological polar surface area (TPSA) using quantitative structure-activity relationship (QSAR) analysis. Several machine learning models, including Linear Regression, Ridge Regression, Lasso Regression, Random Forest Regression, and Gradient Boosting, were developed to improve prediction accuracy using topological indices. The datasets were combined with appropriate topological indices for individual compounds. Model performance was evaluated using Mean Squared Error (MSE) and score after hyperparameter tuning via GridSearchCV. Ridge and Lasso Regression models stood out due to their lowest Test MSE averages (3617.74 and 3540.23, respectively) and highest scores (0.9322 and 0.9374, respectively), demonstrating their effectiveness in handling multicollinearity and preventing overfitting. Linear Regression also performed robustly, achieving an MSE of 5249.97 and an of 0.8563, highlighting the suitability of simpler models for datasets with inherent linear relationships. While Random Forest and Gradient Boosting Regression are capable of capturing nonlinear relationships, their performance varied. Random Forest Regression achieved an MSE of 6485.45 and an of 0.6643, while Gradient Boosting initially performed poorly with an MSE of 4488.04 and an of 0.5659. After fine-tuning Gradient Boosting with an expanded hyperparameter grid, its performance improved significantly, achieving a Test MSE of 1494.74 and an of 0.9171. However, it still ranked fourth, suggesting that simpler models like Linear, Ridge, and Lasso Regression may be better suited for this dataset. This work emphasizes the significance of accurate model selection and optimization in QSAR analysis, demonstrating how these approaches can be used to develop dependable predictive models in computational drug discovery and cheminformatics.
More Related Videos
00:05In Silico Modeling Method for Computational Aquatic Toxicology of Endocrine Disruptors: A Software-Based Approach Using QSAR Toolbox
Published on: August 28, 2019
10:21Author Spotlight: Streamlining Protein Target Prediction and Validation via Molecular Docking and CETSA
Published on: February 23, 2024
Related Concept Videos
Structure-Activity Relationships and Drug Design
SAR studies the intricate relationship between a drug's chemical structure and biological activity. It focuses on understanding how modifications to a drug's structure can influence...
Model Approaches for Pharmacokinetic Data: Distributed Parameter Models
The distributed parameter models are specifically designed to account for variations and differences in some drug classes. This model is particularly useful for assessing regional concentrations of anticancer or...
Pharmacokinetic Models: Overview
There are three primary types of models: empirical, compartment, and physiological. Empirical models, with minimal...
Quantitative Aspects of Drug-Receptor Interaction
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Analysis Methods of Pharmacokinetic Data: Model and Model-Independent Approaches
The model approach uses mathematical models to describe changes in drug concentration over time. Pharmacokinetic models help characterize drug behavior in patients, predict drug concentration in the body fluids, calculate optimum dosage regimens, and evaluate the risk of toxicity. However, ensuring that the model fits the experimental data accurately...