Related Experiment Videos
DeepFusion-ACSM: an interpretable hybrid ML-DL ensemble model for anticancer small-molecule activity prediction
Deepak Devakumar Sagayaraj1, Priya Dharshini Balaji2, Arvin Anand1
1Department of Computational Intelligence, SRM Institute of Science and Technology, Kattankulathur, Chengalpattu, 603203, Tamil Nadu, India.
Abstract:
Cancer remains one of the leading causes of mortality worldwide, and the discovery of novel anticancer drugs continues to be a costly and time-consuming process. Although machine learning (ML) and deep learning (DL) have significantly advanced computational drug discovery, many existing predictive models rely on a single learning algorithm and often fail to generalize across diverse molecular descriptors, limiting their predictive reliability. In this study, we propose an interpretable hybrid ML-DL ensemble framework for the prediction of anticancer small molecules by integrating complementary learning algorithms, multiple feature selection strategies, and ensemble fusion techniques to improve robustness, generalization, and recall-oriented prediction performance. A total of 102 ML models were developed using 17 algorithms on two molecular descriptor datasets (2D and Hybrid) with three feature selection strategies, while 12 DL models based on Improved Deep MLP, FT-Transformer, and TabNet architectures were optimized using five-fold cross-validation. Out-of-fold predictions from the optimized ML and DL models were further integrated through eight ensemble strategies to construct ML, DL, and hybrid ML-DL ensembles. The best-performing hybrid ensemble, comprising XGBoost, Improved Deep MLP, and TabNet trained on the 2D dataset with mutual information-selected features, achieved a macro-recall of 0.83, an accuracy of 0.83, and a ROC-AUC of 0.90, outperforming the individual ML and DL models. Our developed model correctly identified 8 out of 10 FDA-approved compounds with known anticancer and non-anticancer labels during external validation. Additionally, the framework was applied to 1200 FDA-approved drugs from the ZINC database for prospective virtual screening, where compounds were ranked according to their ensemble prediction probabilities to prioritize. These findings indicate that integrating complementary ML and DL models with ensemble decision fusion enables more accurate and reliable identification of anticancer small molecules than individual predictive models. Overall, the proposed framework provides a reproducible, interpretable, and scalable computational strategy for accelerating early-stage anticancer drug discovery.