Related Experiment Video
Updated: May 27, 2025

Selecting Multiple Biomarker Subsets with Similarly Effective Binary Classification Performances
Published on: October 11, 2018
Feature Selection for Hypertension Risk Prediction Using XGBoost on Single Nucleotide Polymorphism Data
Lailil Muflikhah1, Tirana Noor Fatyanosa1, Nashi Widodo2
1Department of Informatics Engineering, Faculty of Computer Science, Brawijaya University, Malang, Indonesia.
Insights
This study developed an XGBoost model to identify 292 single nucleotide polymorphisms (SNPs) as hypertension biomarkers. The model achieved 98% accuracy, offering a powerful new tool for predicting hypertension risk.
Area of Science:
- Genetics
- Bioinformatics
- Computational Biology
Background:
- Hypertension is a global health concern with severe complications if untreated.
- Identifying reliable biomarkers for hypertension risk is crucial for early detection and prevention.
- Genetic variations, specifically single nucleotide polymorphisms (SNPs), offer potential markers for disease predisposition.
Purpose of the Study:
- To develop a feature selection model using the XGBoost algorithm.
- To identify specific single nucleotide polymorphisms (SNPs) as effective biomarkers for hypertension risk detection.
- To evaluate the performance of the XGBoost model in predicting hypertension.
Main Methods:
- Utilized the OpenSNP dataset comprising 19,697 SNPs from 2,052 samples.
- Employed Extreme Gradient Boosting (XGBoost), an ensemble machine learning method, for feature selection.
- Built a classifier model using high-dimensional genetic variation data (SNPs) for prediction.
Main Results:
- Identified 292 significant SNPs for hypertension risk prediction.
- Achieved high performance metrics: 98.55% F1-score, 98.73% precision, 98.38% recall, and 98% overall accuracy.
- Demonstrated superior performance of XGBoost feature selection compared to genetic algorithms, ANOVA, chi-square, and PCA.
Conclusions:
- Successfully developed a predictive model for hypertension using SNP data.
- Effectively managed high-dimensional SNP data to pinpoint significant features as biomarkers.
- The XGBoost feature selection method shows high efficacy in predicting hypertension risk.
Objectives:
Hypertension, commonly known as high blood pressure, is a prevalent and serious condition affecting a significant portion of the adult population globally. It is a chronic medical issue that, if left unaddressed, can lead to severe health complications, including kidney problems, heart disease, and stroke. This study aims to develop a feature selection model using the XGBoost algorithm to identify specific single nucleotide polymorphisms (SNPs) as biomarkers for detecting hypertension risk.
Methods:
We propose using the high dimensionality of genetic variations (i.e., SNPs) to build a classifier model for prediction. In this study, SNPs were used as markers for hypertension in patients. We utilized the OpenSNP dataset, which includes 19,697 SNPs from 2,052 samples. Extreme gradient boosting (XGBoost) is an ensemble machine learning method employed here for feature selection, which incrementally adjusts weights in a series of steps.
Results:
The experimental results identified 292 SNPs that exhibited high performance, with an F1-score of 98.55%, precision of 98.73%, recall of 98.38%, and overall accuracy of 98%. This study provides compelling evidence that the XGBoost feature selection method outperforms other representative feature selection methods, such as genetic algorithms, analysis of variance, chi-square, and principal component analysis, in predicting hypertension risk, demonstrating its effectiveness.
Conclusions:
We developed a model for predicting hypertension using the SNPs dataset. The high dimensionality of SNP data was effectively managed to identify significant features as biomarkers using the XGBoost feature selection method. The results indicate high performance in predicting the risk of hypertension.
Related Concept Videos
Single Nucleotide Polymorphisms-SNPs
Pleiotropy
Genome-wide Association Studies-GWAS
GWAS does not require the identification of the target gene involved in...
Comparing Copy Number Variations and SNPs
Copy number variations or CNVs are the structural variations that cover more than 1kb of DNA sequence. The single nucleotide polymorphism (SNP), on the other hand, is a single nucleotide change or a point mutation that is found in more than 1%...
Polygenic Traits
Human Genetics
The complex relationship between genetics and psychology is observable through common biological components such...

