Related Experiment Video
Updated: Jul 2, 2026

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
A comparison of four feature selection algorithms applied to multiply-imputed proteomic data.
John H Holmes1, Malek Kamoun, Ajay Israni
1University of Pennsylvania School of Medicine, Philadelphia, PA, USA.
AMIA ... Annual Symposium Proceedings. AMIA Symposium
|August 13, 2008
Summary
Single imputation for missing data can be inaccurate during feature selection. Multiple imputation methods improve knowledge discovery by providing more reliable feature selection in data mining.
Area of Science:
- Data Science
- Statistical Analysis
- Machine Learning
Background:
- Missing data are a common challenge in data analysis.
- Single imputation methods can introduce bias, especially during feature selection.
- Inaccurate feature selection can hinder knowledge discovery.
Purpose of the Study:
- To evaluate the impact of imputation methods on feature selection.
- To highlight the benefits of multiple imputation for data mining.
- To underscore the importance of robust methods for knowledge discovery.
Main Methods:
- Applied feature selection procedures to multiply imputed data.
- Compared results from single imputation versus multiple imputation.
- Analyzed the phenomenon of imputation impact on feature selection.
Main Results:
- Demonstrated that single imputation can lead to inaccurate feature selection.
- Showcased that feature selection on multiply imputed data is more reliable.
- Confirmed the phenomenon of imputation's influence on mining data.
Conclusions:
- Multiple imputation is a crucial technique for accurate feature selection.
- Employing multiple imputation enhances the reliability of knowledge discovery.
- Data mining benefits significantly from advanced imputation strategies.
