Related Experiment Video
Updated: Jul 20, 2025

In Silico Clinical Trials for Cardiovascular Disease
Published on: May 27, 2022
Hyperparameter optimization for cardiovascular disease data-driven prognostic system
Jayson Saputra1, Cindy Lawrencya2, Jecky Mitra Saini2
1Industrial Engineering Department, BINUS Graduate Program - Master of Industrial Engineering, Bina Nusantara University, Jakarta 11480, Indonesia. jayson@binus.ac.id.
Insights
Predicting cardiovascular diseases (CVDs) is crucial. This study used machine learning models, finding Stochastic Gradient Descent (SGD) and Artificial Neural Networks (ANN) achieved high accuracy in CVD risk prediction.
Area of Science:
- Cardiovascular disease research
- Medical data mining
- Machine learning in healthcare
Background:
- Cardiovascular diseases (CVDs) are a leading cause of global mortality, necessitating improved prediction and diagnostic methods.
- Timely prognosis, considering patient history and lifestyle, is key for CVD prevention and management.
- Utilizing patient datasets for CVD risk factor analysis presents a significant challenge in modern medicine.
Purpose of the Study:
- To apply data mining and unsupervised machine learning techniques for analyzing Cardiovascular Disease Prognostic datasets.
- To evaluate the performance and classification accuracy of various machine learning models for CVD prediction.
- To determine the optimal number of clusters within CVD patient data using clustering methods.
Main Methods:
- Employed data mining via Orange software on a dataset of 918 adult patients (28-77 years old).
- Utilized supervised learning algorithms including k-nearest neighbors, support vector machine, random forest, artificial neural network (ANN), naïve bayes, logistic regression, stochastic gradient descent (SGD), and AdaBoost.
- Applied unsupervised clustering methods such as k-means, hierarchical, and density-based spatial clustering of applications with noise (DBSCAN) to identify data patterns.
Main Results:
- Stochastic Gradient Descent (SGD) and Artificial Neural Network (ANN) models demonstrated the highest performance, achieving a classification accuracy of 0.900.
- K-means and hierarchical clustering methods indicated that the Cardiovascular Disease Prognostic datasets could be effectively divided into two distinct clusters.
- The study highlights the strong correlation between model accuracy and the effectiveness of CVD risk prediction.
Conclusions:
- SGD and ANN are highly effective models for predicting cardiovascular disease risk with significant accuracy.
- The clustering of patient data into two groups suggests potential for distinct risk stratification.
- Accurate predictive models are essential for improving diagnostic capabilities and enabling timely preventive interventions for CVD.
Abstract:
Prediction and diagnosis of cardiovascular diseases (CVDs) based, among other things, on medical examinations and patient symptoms are the biggest challenges in medicine. About 17.9 million people die from CVDs annually, accounting for 31% of all deaths worldwide. With a timely prognosis and thorough consideration of the patient's medical history and lifestyle, it is possible to predict CVDs and take preventive measures to eliminate or control this life-threatening disease. In this study, we used various patient datasets from a major hospital in the United States as prognostic factors for CVD. The data was obtained by monitoring a total of 918 patients whose criteria for adults were 28-77 years old. In this study, we present a data mining modeling approach to analyze the performance, classification accuracy and number of clusters on Cardiovascular Disease Prognostic datasets in unsupervised machine learning (ML) using the Orange data mining software. Various techniques are then used to classify the model parameters, such as k-nearest neighbors, support vector machine, random forest, artificial neural network (ANN), naïve bayes, logistic regression, stochastic gradient descent (SGD), and AdaBoost. To determine the number of clusters, various unsupervised ML clustering methods were used, such as k-means, hierarchical, and density-based spatial clustering of applications with noise clustering. The results showed that the best model performance analysis and classification accuracy were SGD and ANN, both of which had a high score of 0.900 on Cardiovascular Disease Prognostic datasets. Based on the results of most clustering methods, such as k-means and hierarchical clustering, Cardiovascular Disease Prognostic datasets can be divided into two clusters. The prognostic accuracy of CVD depends on the accuracy of the proposed model in determining the diagnostic model. The more accurate the model, the better it can predict which patients are at risk for CVD.
Related Concept Videos
Blood Studies for Cardiovascular System I: Cardiac Biomarkers
The essential diagnostic tools for detecting myocardial necrosis and monitoring individuals suspected of having acute coronary syndrome (ACS) include:
Troponins
Troponins, particularly cardiac troponins I and T, are the most precise and sensitive markers of myocardial injury. They are detectable within 4-6 hours of myocardial injury and remain...
Blood Studies for Cardiovascular System II: CRP, Hcy, and Cardiac Natriuretic Peptide Markers
These markers indicate stress or strain on the heart muscle:
Natriuretic Peptides (BNP)
Cardiac myocytes produce these hormones in response to ventricular stretching...
Model Approaches for Pharmacokinetic Data: Physiological Models
Cancer Survival Analysis
Cardiomyopathy V: Interprofessional Care

