Related Experiment Videos
Machine learning-based identification of heart failure with reduced ejection fraction using routine laboratory
Zhiping Meng1, Xuezhan Qin1, Binbin Liang2
1Department of Laboratory Medicine, Guigang Clinical Medical Research Center for Medical Laboratory, Eighth Affiliated Hospital of Guangxi Medical University, Guigang City People's Hospital, Guigang, Guangxi, China.
Background:
Early recognition of heart failure with reduced ejection fraction (HFrEF) is clinically important because treatment pathways and prognosis differ across left ventricular ejection fraction phenotypes. When echocardiography is delayed, routinely available laboratory data may help prioritize patients for definitive imaging.
Methods:
This single-center retrospective study included 1,480 hospitalized patients with chronic heart failure, comprising 377 with HFrEF and 1,103 with HFmrEF/HFpEF. Baseline laboratory indicators were obtained from the first blood draw within 24 h of admission, and matched echocardiography was typically performed within 48 h. After a stratified 7:3 train-test split, preprocessing, imputation, standardization, feature selection, and model training were conducted within a leakage-free machine-learning pipeline. Cross-validated recursive feature elimination retained 13 routine indicators: proBNP, HCT, TBil, DBil, BUN, β2-MG, CRP, UA, Glb, GGT, MCH, LDL-C, and FIB. Six algorithms were evaluated using discrimination, calibration, decision-curve analysis, threshold analysis, and SHAP interpretation.
Results:
The HFrEF group showed higher proBNP, BUN, TBil, DBil, GGT, HCT, and UA levels than the HFmrEF/HFpEF group. In the independent test set, random forest and XGBoost achieved the highest AUCs (both 0.789), whereas logistic regression showed comparable discrimination (AUC: 0.784) and strong calibration (intercept 0.017; slope 1.069). Random forest had the lowest Brier score (0.152), but pairwise bootstrap comparisons showed no statistically significant AUC differences among models. At the default threshold, sensitivity was limited; however, moving the random forest threshold to 0.15 increased sensitivity to 0.912 and negative predictive value to 0.935, supporting a rule-out triage strategy for prioritized echocardiography. SHAP analysis identified proBNP as the dominant predictor, followed by HCT, DBil, Glb, UA, GGT, TBil, and LDL-C.
Conclusions:
A 13-indicator routine laboratory model provided moderate discrimination for distinguishing HFrEF from HFmrEF/HFpEF. Random forest had the numerically most favorable balance of discrimination and calibration, but its clinical usefulness depends on threshold optimization rather than default classification. The model should be regarded as an adjunctive triage tool pending external and prospective validation, not as a substitute for echocardiographic phenotyping.