Leveraging Clinical Data for Early Heart Disease Prediction: A Machine Learning Approach With Interpretability
Emma Qumsiyeh1, Qassam Al-Wirdian1, Nur Sebnem Ersoz2
1Faculty of Engineering and Information Technology, Palestine Ahliya University, Bethlehem, Palestine.
Insights
Machine learning models, particularly Random Forest and K-Nearest Neighbors (KNN), show strong potential for accurate heart disease prediction. Explainable AI (SHAP) enhances model interpretability for clinical decision support.
Area of Science:
- Cardiology
- Medical Informatics
- Artificial Intelligence
Background:
- Heart disease is a major global cause of mortality.
- Early and accurate diagnosis is crucial for effective prevention and treatment.
- This necessitates advanced diagnostic tools for risk stratification.
Purpose of the Study:
- To develop and evaluate machine learning models for heart disease prediction.
- To compare the performance of Logistic Regression, Random Forest, KNN, and Decision Trees.
- To enhance model interpretability using SHAP values for clinical trust.
Main Methods:
- Utilized a publicly available clinical and demographic dataset.
- Performed data preprocessing including imputation, encoding, and normalization.
- Evaluated four classification algorithms (Logistic Regression, Random Forest, KNN, Decision Trees) using accuracy, precision, recall, and AUC-ROC metrics.
- Applied SHapley Additive exPlanations (SHAP) for model interpretability.
Main Results:
- Hyperparameter-optimized Random Forest and KNN models demonstrated superior predictive performance.
- SHAP analysis provided insights into feature importance and individual prediction explanations.
- The study confirmed the effectiveness of machine learning in predicting heart disease.
Conclusions:
- Interpretable machine learning models offer significant potential for early heart disease diagnosis and clinical decision support.
- SHAP enhances transparency and clinical trust in AI-driven diagnostic tools.
- Future work should focus on larger datasets and real-time applications to improve generalizability and clinical utility.
Background:
Heart disease remains one of the leading causes of mortality worldwide, highlighting the need for early and accurate diagnosis to support effective prevention and treatment strategies.
Methods:
This study presents a machine-learning-based approach for predicting heart disease using clinical and demographic data from a publicly available dataset. Four widely used classification algorithms-Logistic Regression, Random Forest, K-Nearest Neighbors (KNN), and Decision Trees-were evaluated to identify the most effective predictive model. The dataset underwent comprehensive preprocessing, including handling missing values, categorical encoding, and feature normalization, to enhance data quality and model robustness. Model performance was assessed using accuracy, precision, recall, and AUC-ROC metrics.
Results:
Findings show that hyperparameter-optimized models, particularly Random Forest and KNN, demonstrated strong predictive performance. Explainability techniques, specifically SHapley Additive exPlanations (SHAP), were incorporated to improve interpretability, transparency, and clinical trust. SHAP values were used to analyze feature importance and provide explanations for individual predictions.
Conclusion:
The results underscore the potential of interpretable machine-learning models as valuable tools for early diagnosis, risk stratification, and clinical decision support. Future research should employ larger datasets and investigate real-time predictive applications further to enhance the generalizability and clinical utility of these models.

