ML-BUSMetab: Machine Learning-Based Metabolomic Profiling for Predicting Aspirin Response in Colorectal Cancer
Abdulvahap Pınar1,2, Ahmet Kadir Arslan1, Cemil Çolak1
1Department of Biostatistics and Medical Informatics, Faculty of Medicine, İnönü University, 44280 Malatya, Turkey.
None:
Background/Objectives: Aspirin-based colorectal cancer (CRC) chemoprevention remains a promising yet individually variable strategy. As a proof-of-concept toward future personalized chemoprevention frameworks, we aimed to develop and validate machine learning (ML) models capable of distinguishing aspirin-exposed from placebo-exposed participants based on their plasma metabolomic signatures, thereby characterizing the metabolomic footprint of aspirin administration rather than directly predicting clinical chemoprevention benefit. Methods: Training was performed on the Aspirin/Folate Polyp Prevention Study (AFPPS) dataset ST001422 (n = 300) and external validation on ST001423 (n = 223). After multi-method consensus feature selection, reducing 19,433 features to 300, sixteen ML and deep learning (DL) architectures were benchmarked under nested cross-validation. Model interpretability was assessed using SHapley Additive exPlanations (SHAP) and Local Interpretable Model-agnostic Explanations (LIME) analyses. Results: GBM_sklearn achieved the highest cross-validation Precision-Recall AUC (PR-AUC) of 0.945, while ensemble stacking (Stack_LGB) offered superior calibration (Brier = 0.117). DL models consistently underperformed traditional ML (PR-AUC: 0.673-0.843 vs. 0.881-0.945), attributable to limited sample size. SHAP and LIME analyses independently identified m/z 196.0604 (C18, RT 89.4 s) as the top metabolic biomarker, consistent with aspirin-induced glycerophospholipid pathway alterations. External validation performance degraded substantially (PR-AUC: 0.945 → 0.711), attributable to inter-study analytical batch effects. Conclusions: This framework demonstrates the feasibility of metabolomics-driven personalized chemoprevention. Although the high feature-to-sample ratio (300:300) and the substantial drop between internal and external performance indicate that the cross-validation estimates likely include dataset-specific noise in addition to the true biological signal. While highlighting batch harmonization and aggressive feature reduction (e.g., LASSO/RFE-based selection of 10-20 high-impact metabolites) as a prerequisite for clinical translation.

