Related Experiment Video
Updated: Sep 16, 2026

Detection of Residual Donor Erythroid Progenitor Cells after Hematopoietic Stem Cell Transplantation for Patients with Hemoglobinopathies
Published on: September 6, 2017
Development and Internal Validation of a Machine Learning-Based Model for Thalassemia Identification Using Complete
Ping Yin1, Jialin Tan1, Honghai Hong1
1Department of Clinical Laboratory, Guangdong Provincial Key Laboratory of Major Obstetric Diseases, Guangdong Provincial Clinical Research Center for Obstetric and Gynecology, The Third Affiliated Hospital, Guangzhou Medical University, Guangzhou 510150, China.
Abstract:
Background/Objectives: While genetic testing for thalassemia is definitive yet costly for population screening, and conventional methods lack specificity, this study aims to develop and internally validate a machine learning (ML) model based on routine complete blood count (CBC) parameters to identify individuals at increased risk of thalassemia who may benefit from confirmatory genetic testing. Methods: Data from 1820 individuals were retrospectively collected and randomly divided using stratified sampling into training (70%) and internal validation (30%) cohorts. Feature selection was conducted within the training cohort, reducing 225 analyzer-derived parameters to 12 features. SMOTE was applied to the training portion of each five-fold cross-validation split. Seven ML models were evaluated in the validation cohort, with post hoc global feature attribution assessed using SHAP. Results: Seven ML models were developed using 12 selected features. All performed strongly (AUC: 0.885-0.918), with Random Forest (RF) achieving the highest AUC of 0.918 (95% CI: 0.900-0.940), accuracy of 85.7%, sensitivity of 89.7%, specificity of 83.0%, and F1 score of 0.837, with a PPV of 78.4% and NPV of 92.1%. SHAP analysis identified mean corpuscular volume (MCV), mean corpuscular hemoglobin (MCH), and red blood cell distribution width-standard deviation (RDW-SD) as the most influential features. Conclusions: The RF-based model showed good discrimination in validation, while SHAP analysis provided information on global feature contributions. The model may assist in identifying individuals who warrant confirmatory testing; however, prospective external validation is required before routine clinical implementation.

