Machine learning for discovering anticancer compounds from plants and elucidating their mechanisms
Chenyu Zhang1, Shanxue Jiang2,3, Yushuang Li4
1School of Light Industry Science and Engineering, Beijing Technology and Business University, Beijing, 102488, China.
Abstract:
Plants are important sources of bioactive compounds with demonstrated anticancer and cancer-preventive properties. However, traditional methods for discovering and deciphering the mechanisms of these compounds are often slow and labor-intensive. The integration of machine learning (ML) provides a computational approach for bioactivity prediction, candidate prioritization, and mechanistic hypothesis generation. Algorithms including support vector machines (SVM) and graph neural networks (GNN) are increasingly employed to support multiple stages from compound screening to experimentally guided mechanistic investigation. Despite these advances, challenges such as data heterogeneity and structural complexity of natural products remain. This review systematically outlines the natural anticancer compounds derived from plants, summarizes the application of machine learning in bioactivity prediction, candidate prioritization, and mechanistic investigation, compares functional-group prevalence and co-occurrence patterns in plant-derived anticancer compounds and FDA-approved anticancer drugs, and identifies current challenges along with future directions for AI-assisted discovery in this field.
