Optimizing Stroke Risk Prediction: A Primary Dataset-Driven Ensemble Classifier With Explainable Artificial
Md Maruf Hossain1,2, Md Mahfuz Ahmed1,2, Md Rakibul Hasan Rakib1
1Department of Biomedical Engineering Islamic University Kushtia Bangladesh.
Background And Aims:
Stroke remains a leading cause of mortality and long-term disability worldwide, presenting a significant global health challenge. Effective early prediction models are essential for reducing its impact. This study introduces a novel ensemble method for predicting stroke using two datasets: a primary dataset collected from a hospital, containing medical histories and clinical parameters, and a secondary dataset.
Methods:
We applied several preprocessing techniques, including outlier detection, data normalization, k-means clustering, and missing value detection, to refine the datasets. A novel ensemble classifier was developed, combining AdaBoost, Gradient Boosting Machine (GBM), Multilayer Perceptron (MLP), and Random Forest (RF) algorithms to enhance predictive accuracy. Additionally, Explainable Artificial Intelligence (XAI) techniques such as SHAP and LIME were integrated to elucidate key features influencing stroke prediction.
Results:
The proposed ensemble classifier achieved an accuracy of 95% for the secondary dataset and 80.36% for the primary dataset. Comparative analysis with other machine learning models highlighted the superior performance of the ensemble approach. The integration of XAI further provided insights into the critical indicators influencing stroke classification, improving model interpretability and decision-making.
Conclusion:
Our study demonstrates that the novel ensemble classifier, supported by effective preprocessing and XAI techniques, is a powerful tool for stroke prediction. The high accuracy rates achieved validate its effectiveness and potential for practical clinical application. Future work will focus on incorporating deep learning techniques and medical imaging to further improve classification accuracy and model performance.
Related Concept Videos
Prediction Intervals
However, the point estimate is most likely not the exact value of the population parameter, but close to it. After calculating point estimates, we construct interval estimates, called confidence intervals or prediction intervals. This prediction interval comprises a range of values unlike the point estimate and is a better predictor of the observed sample value, y.
Sensitivity, Specificity, and Predicted Value
Sensitivity is the...
Aggregates Classification
Petrographic classification groups aggregates based on common mineralogical characteristics. Some of the common mineral groups found in aggregates are...
Improving Translational Accuracy
Classification of Systems-I
Homogeneity dictates that if an input x(t) is multiplied by a constant c, the output y(t) is multiplied by the same constant. Mathematically, this is expressed as:
Receiver Operating Characteristic Plot


