Related Experiment Video
Updated: Apr 19, 2026

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Hybrid feature-selection and diversity-guided stacking framework for interpretable ensemble learning: Application to
Farideh Mohtasham1, Seyed Saeed Hashemi Nazari2, Mohamad Amin Pourhoseingholi3
1Gastroenterology and Liver Diseases Research Center, Research Institute for Gastroenterology and Liver Diseases, Shahid Beheshti University of Medical Sciences, Tehran, Iran.
This study introduces a hybrid ensemble learning framework that enhances predictive accuracy and interpretability in high-dimensional data. The novel approach balances model diversity and feature selection for robust and scalable machine learning applications.
Area of Science:
- Biomedical data science
- Machine learning
- Ensemble methods
Background:
- High-dimensional biomedical data presents challenges for predictive modeling, requiring accuracy, interpretability, and efficiency.
- Existing ensemble methods often lack model diversity and use suboptimal feature selection, limiting generalizability.
- A novel hybrid framework is proposed to enhance robustness and scalability in data-intensive domains.
Purpose of the Study:
- To develop a hybrid feature-selection and diversity-guided stacking framework for improved predictive modeling.
- To address limitations in current ensemble methods regarding model diversity and feature selection.
- To enhance the robustness and scalability of machine learning models in clinical and other data-intensive fields.
Main Methods:
- A hybrid feature-selection pipeline combining Variance Inflation Factor (VIF), Analysis of Variance (ANOVA), Sequential Backward Elimination (SBE), and Lasso regression was employed.
- A diversity-aware stacking strategy utilized pairwise (Disagreement, Yule's Q, Cohen's Kappa) and non-pairwise (Entropy, Kohavi-Wolpert) diversity metrics.
- The framework was validated on COVID-19 patient data using robust scaling and ROSE-based class balancing with 16 base classifiers and 5 meta-learners via 10-fold cross-validation.
Main Results:
- The optimal configuration achieved 91.4% accuracy, with an AUC of 0.955, outperforming individual models.
- Computational efficiency was demonstrated with a training time of ~450s and inference time <0.2s per case.
- Feature importance and SHAP analysis confirmed clinical relevance and model interpretability.
Conclusions:
- The proposed framework effectively improves predictive accuracy and interpretability while maintaining computational efficiency.
- The approach is broadly applicable to various prediction tasks in biomedical, environmental, and engineering fields.
- This method offers a scalable and interpretable solution for ensemble learning in complex datasets.
Related Concept Videos
Types of Selection
Multi-input and Multi-variable systems
In the absence of...
Improving Translational Accuracy
Improving Translational Accuracy
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...