Related Experiment Video
Updated: May 5, 2026

A Rapid Screening Workflow to Identify Potential Combination Therapy for GBM using Patient-Derived Glioma Stem Cells
Published on: March 28, 2021
XGBPred-ACSM: A Hybrid Descriptor-Driven XGBoost Framework for Anticancer Small Molecule Prediction
Priya Dharshini Balaji1, Subathra Selvam1, Anuradha Thiagarajan2
1Computational Biology Laboratory, Department of Genetic Engineering, School of Bioengineering, SRM Institute of Science and Technology, Kattankulathur, Chengalpattu 603203, Tamil Nadu, India.
Abstract:
Background/Objectives: Cancer remains one of the leading global health burdens, mainly because of the lack of specificity and off-target toxicity associated with conventional therapeutic approaches. To move toward more efficient anticancer drug discovery, we have developed an advanced machine-learning-based architecture that allows for predictive modeling of anticancer small molecules. Methods: A total of 3600 compounds with experimentally validated IC50 values were systematically processed to derive a comprehensive suite of molecular representations comprising 2D physicochemical descriptors, structural fingerprints, and hybrid descriptor sets generated via the Mordred and PaDEL frameworks. A total of six machine learning algorithms-Random Forest (RF), Extreme Gradient Boosting (XGB), Gradient Boosting (GB), Extra-Trees classifier (ET), Adaptive Boosting (AdaBoost), and Light Gradient Boosting Machine (LightGBM)-were trained and benchmarked via a rigorous model evaluation protocol incorporating 10-fold cross-validation along with multiple performance metrics. Ensemble voting strategies were also examined to assess potential performance. Result: Of all configurations, the XGB-Hybrid architecture emerged as the most robust and generalizable classifier with an AUC of 0.88 and accuracy of 79.11% on the independent test set. To ensure interpretability and mechanistic insight, SHAP-based feature analysis was conducted, by which feature contributions could be quantified and the molecular determinants most influential for anticancer activity discrimination were revealed. Altogether, the current study establishes an XGB-Hybrid framework as technically rigorous, interpretable, and high-performance predictive modeling with the ability to accelerate early-stage anticancer small molecule identification. Conclusions: The study has brought into focus the transformational effect of machine learning in modern computational oncology and rational drug design pipelines.
Insights
Machine learning accelerates anticancer drug discovery by predicting small molecule efficacy. An XGB-Hybrid model achieved 79.11% accuracy, identifying key molecular features for targeted therapies.
Area of Science:
- Computational oncology
- Drug discovery and design
- Machine learning applications in medicine
Background:
- Cancer poses a significant global health challenge due to limitations in current therapeutic specificity and toxicity.
- Developing more effective anticancer drugs requires advanced predictive modeling approaches.
- Machine learning offers a powerful tool for enhancing the efficiency of anticancer drug discovery.
Purpose of the Study:
- To develop and validate an advanced machine learning-based architecture for predicting anticancer small molecules.
- To systematically process molecular representations and benchmark various machine learning algorithms.
- To identify the most influential molecular determinants for anticancer activity.
Main Methods:
- Utilized 3600 compounds with experimentally validated IC50 values.
- Derived molecular representations using 2D physicochemical descriptors, structural fingerprints, and hybrid sets (Mordred, PaDEL).
- Trained and evaluated six machine learning algorithms (RF, XGB, GB, ET, AdaBoost, LightGBM) using 10-fold cross-validation and SHAP-based feature analysis.
Main Results:
- The XGB-Hybrid architecture demonstrated superior performance with an AUC of 0.88 and 79.11% accuracy on an independent test set.
- SHAP analysis provided mechanistic insights by quantifying feature contributions and identifying key molecular determinants.
- The developed framework is technically rigorous, interpretable, and high-performing for early-stage anticancer small molecule identification.
Conclusions:
- Machine learning, particularly the XGB-Hybrid framework, significantly advances computational oncology and rational drug design.
- This approach accelerates the identification of potential anticancer small molecules.
- The study highlights the transformative impact of machine learning in modern drug discovery pipelines.
Related Concept Videos
Adaptive Mechanisms in Cancer Cells
Some of the advantages that cancer cells have on normal cells include - enhanced ability to divide without terminally differentiating, induce new blood vessel formation,...
Cancer Survival Analysis
Combination Therapies and Personalized Medicine
The combination of the drug acetazolamide and sulforaphane is a good example of combination therapy to treat cancer. The cells in the interior of a large tumor often die due to the hypoxic and...

