Hyperparameter optimization for cardiovascular disease data-driven prognostic system

Jayson Saputra1, Cindy Lawrencya2, Jecky Mitra Saini2

  • 1Industrial Engineering Department, BINUS Graduate Program - Master of Industrial Engineering, Bina Nusantara University, Jakarta 11480, Indonesia. jayson@binus.ac.id.

Insights

Predicting cardiovascular diseases (CVDs) is crucial. This study used machine learning models, finding Stochastic Gradient Descent (SGD) and Artificial Neural Networks (ANN) achieved high accuracy in CVD risk prediction.

Area of Science:

  • Cardiovascular disease research
  • Medical data mining
  • Machine learning in healthcare

Background:

  • Cardiovascular diseases (CVDs) are a leading cause of global mortality, necessitating improved prediction and diagnostic methods.
  • Timely prognosis, considering patient history and lifestyle, is key for CVD prevention and management.
  • Utilizing patient datasets for CVD risk factor analysis presents a significant challenge in modern medicine.

Purpose of the Study:

  • To apply data mining and unsupervised machine learning techniques for analyzing Cardiovascular Disease Prognostic datasets.
  • To evaluate the performance and classification accuracy of various machine learning models for CVD prediction.
  • To determine the optimal number of clusters within CVD patient data using clustering methods.

Main Methods:

  • Employed data mining via Orange software on a dataset of 918 adult patients (28-77 years old).
  • Utilized supervised learning algorithms including k-nearest neighbors, support vector machine, random forest, artificial neural network (ANN), naïve bayes, logistic regression, stochastic gradient descent (SGD), and AdaBoost.
  • Applied unsupervised clustering methods such as k-means, hierarchical, and density-based spatial clustering of applications with noise (DBSCAN) to identify data patterns.

Main Results:

  • Stochastic Gradient Descent (SGD) and Artificial Neural Network (ANN) models demonstrated the highest performance, achieving a classification accuracy of 0.900.
  • K-means and hierarchical clustering methods indicated that the Cardiovascular Disease Prognostic datasets could be effectively divided into two distinct clusters.
  • The study highlights the strong correlation between model accuracy and the effectiveness of CVD risk prediction.

Conclusions:

  • SGD and ANN are highly effective models for predicting cardiovascular disease risk with significant accuracy.
  • The clustering of patient data into two groups suggests potential for distinct risk stratification.
  • Accurate predictive models are essential for improving diagnostic capabilities and enabling timely preventive interventions for CVD.

Related Concept Videos

Blood Studies for Cardiovascular System I: Cardiac Biomarkers01:20

Blood Studies for Cardiovascular System I: Cardiac Biomarkers

Cardiac biomarkers are enzymes, proteins, and hormones released into the blood when cardiac cells are injured. They are powerful tools for triaging.
The essential diagnostic tools for detecting myocardial necrosis and monitoring individuals suspected of having acute coronary syndrome (ACS) include:
Troponins
Troponins, particularly cardiac troponins I and T, are the most precise and sensitive markers of myocardial injury. They are detectable within 4-6 hours of myocardial injury and remain...
201
Blood Studies for Cardiovascular System II: CRP, Hcy, and Cardiac Natriuretic Peptide Markers01:19

Blood Studies for Cardiovascular System II: CRP, Hcy, and Cardiac Natriuretic Peptide Markers

Cardiac biomarkers are critical in diagnosing, prognosing, and managing cardiovascular diseases. Routine measurement of specific biomarkers such as B-type natriuretic peptide (BNP), C-reactive protein (CRP), and homocysteine (Hcy) is common practice in clinical settings to evaluate heart function and predict cardiovascular events.
These markers indicate stress or strain on the heart muscle:
Natriuretic Peptides (BNP)
Cardiac myocytes produce these hormones in response to ventricular stretching...
119
Model Approaches for Pharmacokinetic Data: Physiological Models01:15

Model Approaches for Pharmacokinetic Data: Physiological Models

Physiological models in pharmacokinetics are instrumental in understanding the distribution and elimination of drugs within the body. These models describe the drug concentration within target organs, influenced by factors such as drug uptake, tissue volume, and blood flow. Drug uptake is governed by the partition coefficient, which signifies the drug concentration ratio in tissue to that in the blood. The blood flow rate to a specific tissue is expressed as Qt, and the rate of change in tissue...
76
Cancer Survival Analysis01:21

Cancer Survival Analysis

Cancer survival analysis focuses on quantifying and interpreting the time from a key starting point, such as diagnosis or the initiation of treatment, to a specific endpoint, such as remission or death. This analysis provides critical insights into treatment effectiveness and factors that influence patient outcomes, helping to shape clinical decisions and guide prognostic evaluations. A cornerstone of oncology research, survival analysis tackles the challenges of skewed, non-normally...
384
Cardiomyopathy V: Interprofessional Care01:29

Cardiomyopathy V: Interprofessional Care

Managing cardiomyopathy involves addressing underlying or precipitating causes, treating heart failure with medications, and implementing dietary changes and a balanced exercise and rest regimen.Lifestyle ModificationsCardiomyopathy patients should adopt a low-sodium diet to reduce fluid retention and manage heart failure. A personalized exercise and rest plan helps maintain physical fitness without overstraining the heart. Avoiding alcohol and tobacco is essential to prevent further damage to...
16