Related Experiment Video
Updated: May 23, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A hybrid ensemble approach for diabetes prediction using consensus-based feature selection
JianHao Chen1,2, Yung-Wey Chong2, LiLi Wang1,2
1School of Artificial Intelligence, Chongqing Youth Vocational & Technical College, Chongqing, China.
Background:
Diabetes mellitus affects approximately 589 million adults worldwide, with a large proportion remaining undiagnosed until complications arise. Accurate, data-driven early detection tools are urgently needed to support timely clinical intervention.
Objectives:
This study aimed to develop a hybrid ensemble framework integrating multiple feature selection strategies with a consensus approach to improve diabetes prediction accuracy and clinical interpretability.
Methods:
A publicly available dataset of 1,879 patients with 46 features was analysed. Six interaction features (e.g., HbA1c/FBS ratio, Age×BMI) were engineered. Four supervised selection methods, namely Recursive Feature Elimination (RFE), Random Forest, ANOVA, and Mutual Information, were combined with Principal Component Analysis (PCA), and a consensus criterion (≥3 of 4 supervised methods) defined the final feature subset. Seven classifiers were individually optimised via GridSearchCV with 5-fold stratified cross-validation and integrated into a soft voting ensemble, evaluated on a stratified held-out test set (20%, n = 376) using accuracy, precision, recall, F1-score, and AUC-ROC with 95% confidence intervals.
Results:
Eight consensus features were identified, namely HbA1c, fasting blood sugar, hypertension, excessive thirst, frequent urination, cholesterol LDL, BP_diff, and HbA1c/FBS ratio, reducing dimensionality by 82.6% (46 to 8 features). The soft voting ensemble achieved an AUC of 0.948 (95% CI: 0.926-0.970), accuracy of 92.6% (95% CI: 0.890-0.952), precision of 91.8%, recall of 89.4%, and F1-score of 90.6%, outperforming all individual classifiers in recall and F1-score.
Conclusion:
The proposed framework combines supervised and unsupervised feature selection with ensemble learning, yielding a clinically interpretable and high-performing diabetes prediction model. Its consensus-driven feature transparency and robust generalisation support deployment in early screening and digital health applications. Future work should prioritise external multi-centre validation, explainable AI integration, and real-time clinical decision support development.
Related Concept Videos
Diabetic Nephropathy
Diabetes Mellitus: Introduction
Diabetes Mellitus: Overview and Type I Subtype
Type 1 diabetes is an autoimmune disease in which the immune system mistakenly attacks and destroys the insulin-producing beta cells in the pancreas. As a result, the body is unable to produce sufficient insulin, and individuals with...
Diabetes: Management and Pharmacotherapy
Insulin remains the cornerstone of treatment for most patients with type 1 and many...