Related Experiment Video
Updated: Aug 18, 2025

Competing-Risk Nomogram for Predicting Cancer-Specific Survival in Multiple Primary Colorectal Cancer Patients after Surgery
Published on: September 27, 2024
Classification of breast cancer recurrence based on imputed data: a simulation study
Rahibu A Abassi1, Amina S Msengwa2
1Department of Natural Sciences, State University of Zanzibar, Zanzibar, Tanzania. rahibuabassi@yahoo.com.
This study compares statistical classification methods for breast cancer recurrence with missing data. Linear discriminant analysis showed highest accuracy, while logistic regression and k-nearest neighbor performed best for predicting recurrence risk.
Area of Science:
- Medical Statistics
- Machine Learning in Healthcare
- Biostatistics
Background:
- Limited research exists on statistical classification accuracy for breast cancer recurrence, especially with imputed missing data.
- Understanding classifier performance with missing data is crucial for reliable medical event prediction.
Purpose of the Study:
- To compare the performance of binary classifiers (logistic regression, linear discriminant analysis, quadratic discriminant analysis) on breast cancer recurrence.
- To evaluate how imputed missing data affects classifier accuracy and discriminative ability.
Main Methods:
- Simulated incomplete datasets with 15-60% missingness under Missing At Random (MAR) and Missing Completely At Random (MCAR) mechanisms.
- Employed various imputation techniques: mean, hot deck, k-nearest neighbor, multiple imputations via chained equation, expectation-maximization, and predictive mean matching.
- Compared classification accuracy and Area Under the Receiver Operating Characteristic (ROC) curves for each classifier.
Main Results:
- Linear discriminant analysis achieved the highest classification accuracy (73.9%) using mean imputation with 45% missing data under MCAR.
- Logistic regression with predictive mean matching imputation yielded the highest Area Under ROC curves (0.6418) at 30% missingness (MCAR).
- K-nearest neighbor imputation achieved the highest Area Under ROC curves (0.6428) at 60% missing data under MCAR.
Conclusions:
- The choice of imputation method significantly impacts the performance of statistical classifiers for breast cancer recurrence.
- Different classifiers exhibit varying robustness to missing data, highlighting the need for careful selection in medical applications.
- Findings provide insights into optimizing statistical models for predicting breast cancer recurrence in the presence of missing data.
Related Concept Videos
Cancer Survival Analysis
Comparing the Survival Analysis of Two or More Groups
Kaplan-Meier Approach
Actuarial Approach
Consider the example of a high-risk surgical procedure with significant early-stage mortality. A two-year clinical study is conducted,...

