Related Experiment Videos
Explainable machine learning for hypertension prevalence classification: a cross-sectional study in Hainan Province,
Yuewei Wu1, Shan Huang1, Miaomiao Qi1
1Key Laboratory of Emergency and Trauma of Ministry of Education, Department of Cardiology, The First Affiliated Hospital, Hainan Medical University, Haikou, China.
Objective:
This study used explainable machine learning models to classify prevalent hypertension status in the general population.
Methods:
This cross-sectional study used clinical data collected with questionnaires, physical examinations, blood biochemistry, and routine urine tests from 4,800 permanent residents aged ≥18 years old in Hainan Province, China, from 2021 to 2022. The random forest algorithm was applied to select the most significant features based on importance scores of all variable features. Models for hypertension prevalence classification were created using six machine learning techniques. These models were then compared to select a model with high classification performance based on classification accuracy and the area under curve (AUC) values. Calibration assessment and decision curve analysis were additionally performed to assess clinical applicability. To evaluate and illustrate the best models, the SHapley Additive Explanation (SHAP) values and the Local Interpretable Model-Agnostic Explanations (LIME) algorithms were used.
Results:
In total, 4,606 permanent residents were included in this study (hypertension, 32.5%). They were randomly split into two groups: a training set (3,224, 70%) and a validation set (1,382, 30%). After the random forest algorithm was applied to score feature importance, the top 10 most important features, including age, smoking, urine microalbumin (UALB), educational level, diabetes, body mass index (BMI), sex, triglyceride (TG), income level, and family history of hypertension, were finally included for model construction. With an AUC of 0.8461, the eXtreme Gradient Boosting (XGBoost) model had the greatest performance. The SHAP values were used to quantify the contribution of each input feature to the XGBoost model and demonstrate the importance ranking of predictors. The LIME algorithm integrated the SHAP values to provide a more compelling explanation for each individual classification result. A web page based on the results of this study was developed to support hypertension prevalence screening in clinical environments.
Conclusion:
A machine learning model based on clinical variables was developed and verified, which showed superior performance in identifying prevalent hypertension cases in the general population, opening up new possibilities for the rapid population screening and prevention of hypertension.