Related Experiment Video
Updated: Aug 23, 2025

07:35
Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
7.6K
A Review of Feature Selection Methods for Machine Learning-Based Disease Risk Prediction
Nicholas Pudjihartono1, Tayaza Fadason1,2, Andreas W Kempa-Liehr3
1Liggins Institute, University of Auckland, Auckland, New Zealand.
Frontiers in Bioinformatics
|October 28, 2022
Summary
Feature selection enhances machine learning models for disease risk prediction using genetic data. This approach identifies informative single nucleotide polymorphisms (SNPs) to overcome the curse of dimensionality in precision medicine.
Area of Science:
- Genomics
- Bioinformatics
- Computational Biology
Background:
- Machine learning (ML) excels at pattern detection in complex datasets.
- ML applications in precision medicine aim to predict disease risk using patient genetic data.
- Genotype data presents challenges like the curse of dimensionality (more features than samples).
Purpose of the Study:
- To provide an overview of feature selection methods for ML models.
- To highlight the importance of feature selection in disease risk prediction.
- To focus on identifying relevant single nucleotide polymorphisms (SNPs) for predictive models.
Main Methods:
- Review of various feature selection techniques.
- Analysis of advantages and disadvantages of different methods.
- Discussion of use cases for feature selection in genetic data analysis.
Main Results:
- Feature selection is crucial for improving ML model generalizability.
- Identifying informative SNPs enhances the accuracy of disease risk prediction.
- Methods aim to remove non-informative, irrelevant, and redundant features.
Conclusions:
- Feature selection is vital for robust ML-based disease risk prediction.
- Effective feature selection can mitigate the curse of dimensionality in genomic data.
- This overview aids in selecting appropriate methods for SNP-based risk prediction.

