Related Experiment Video
Updated: Aug 13, 2026

Machine Learning Algorithms for Early Detection of Bone Metastases in an Experimental Rat Model
Published on: August 16, 2020
Exploring prognostic factors in breast cancer: development and selection of optimal machine learning models
Meiying Shen1, Yulei Wang2, Zongyuan Wu1
1Department of Breast Surgery, Maoming People's Hospital, Maoming, China.
Objective:
To investigate prognostic factors for breast cancer recurrence and metastasis, and to systematically compare multiple machine learning models to develop an optimal predictive tool.
Methods:
We retrospectively analyzed data from 1,056 breast cancer patients diagnosed between January 2012 and June 2021 at a single center. Patients were randomly divided into a training (n=740) and a validation (n=316) set. Univariate and multivariate Cox proportional hazards regression analyses were performed to identify independent prognostic factors. Seven machine learning algorithms (Cox regression, LASSO, Elastic-Net, Decision Tree, Random Forest, XGBoost, and GBM) were employed. All models were implemented using survival-specific adaptations. A rigorous 5-fold cross-validation framework was used for model training and hyperparameter tuning. Model performance was evaluated using time-dependent Area Under the Curve (AUC), Brier scores, calibration curves, calibration-in-the-large, calibration slopes, and Decision Curve Analysis (DCA). SHAP values were employed for model interpretation.
Results:
Multivariate Cox regression revealed that tumor size (cm) (HR = 1.025, 95%CI: 1.013-1.037), lymph node dissection (HR = 0.278, 95%CI: 0.199-0.389), ER% (HR = 1.006, 95%CI: 1.003-1.009), PR% (HR = 1.005, 95%CI: 1.002-1.008), Ki-67% (HR = 1.012, 95%CI: 1.007-1.016), and HER2 status (HR = 1.195, 95%CI: 1.098-1.301) were independently associated with disease-free survival. Random Forest and XGBoost demonstrated superior and stable predictive performance, with Random Forest achieving time-dependent AUCs of 0.851 (95%CI: 0.802-0.900, 1-year), 0.763 (95%CI: 0.702-0.824, 3-year), and 0.826 (95%CI: 0.766-0.886, 5-year) in the validation set. The Brier scores for Random Forest were consistently low, and calibration metrics confirmed excellent calibration. DCA indicated a positive net benefit across a wide range of threshold probabilities.
Conclusion:
This single-center study identifies key prognostic factors for breast cancer and demonstrates that ensemble machine learning models, particularly Random Forest, offer superior predictive power. The integration of SHAP interpretation provides a methodological framework for potential clinical application. However, all predictive models require external, multi-center validation before clinical consideration. These findings provide a promising methodological basis, but caution is warranted against overinterpretation until independent verification is complete.

