Related Experiment Videos
From Machine Learning-Enhanced Proteomics to a Validated Diagnostic Model: A Pipeline for Breast Cancer Biomarker
Xiaoyan Zhou1, Yue Li1, Ting Ding1
1Department of Clinical Laboratory, Second Affiliated Hospital of Xi'an Jiaotong University, Xi'an 710004, China.
Abstract:
Early diagnosis of breast cancer (BC) remains challenging. The limited sensitivity and specificity of existing serum tumor markers for reliable clinical application highlight the need to develop a more accurate and efficient screening workflow. This study analyzed serum samples from 255 breast cancer patients and 300 healthy controls using matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry, identifying 58 differentially expressed peptides (37 upregulated, 21 downregulated). Combined with machine learning, peptide identification, and external validation, a complete standardized workflow was established. Nine machine learning (ML) algorithms were employed and compared, including SVM, LightGBM, XGBoost, etc. The models were interpreted using SHAP and LIME to identify key features. Peptides of interest were sequenced via mass spectrometry. Their expression and potential prognostic value were further validated in breast cancer transcriptomic datasets. Nine machine learning algorithms showed favorable discriminatory ability in the study cohort. The LightGBM model achieved an AUC of 0.97 internally and maintained an AUC of 0.88, an accuracy of 0.8543, and a precision of 0.9799 externally. However, after correcting for the markedly elevated prevalence (80.3%) in the external cohort, the positive predictive value (PPV) decreased substantially under real-world screening scenarios, warranting prospective validation in true screening populations. Model interpretation and subsequent sequencing identified six core biomarker peptides: Apolipoprotein A-IV (APOA4), Serum Deprivation Response Protein (SDPR), Alpha-1-Antitrypsin (SERPINA1), Ezrin (EZR), Serglycin (SRGN), and Fibrinogen Alpha Chain (FGA). Transcriptomic corroboration suggested that these molecules were significantly dysregulated in breast cancer tissues and showed univariate prognostic associations with patient survival. These findings demonstrated the potential of a proteomics-driven integrated machine learning pipeline as a proof-of-concept auxiliary risk-stratification tool for enhancing early breast cancer diagnosis, warranting further prospective validation in real-world screening cohorts before clinical translation.