Related Experiment Video
Updated: Jun 11, 2025

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
Construction of machine learning diagnostic models for cardiovascular pan-disease based on blood routine and
Zhicheng Wang1,2,3, Ying Gu1, Lindan Huang1,2
1Institute for Clinical Medical Research, School of Medicine, The First Affiliated Hospital of Xiamen University, Xiamen University, Xiamen, 361003, Fujian, China.
Insights
Machine learning models accurately diagnose cardiovascular diseases using blood test data. Key indicators like potassium and albumin help identify these conditions, paving the way for cost-effective diagnosis and prevention.
Area of Science:
- Biomedical Informatics
- Computational Biology
- Cardiology
Background:
- Cardiovascular disease is a leading global cause of death.
- Traditional diagnostic methods are often costly and time-consuming.
- Accessible blood data offers a potential alternative for diagnosis.
Purpose of the Study:
- To develop machine learning models for cardiovascular disease diagnosis using blood routine and biochemical data.
- To identify unique hematologic and metabolic features associated with cardiovascular diseases.
Main Methods:
- Utilized a dataset of 25,794 healthy individuals and 32,822 cardiovascular disease patients.
- Constructed models using logistic regression, random forest, support vector machine, XGBoost, and deep neural networks.
- Employed the SHAP algorithm for model interpretation.
Main Results:
- The XGBoost model achieved high performance (AUC: 0.9921) for cardiovascular disease prediction.
- Biochemical markers such as potassium, total protein, albumin, and indirect bilirubin were key predictors.
- Specific features like red blood cell count and glucose differentiated various cardiovascular conditions.
Conclusions:
- Machine learning models effectively diagnose cardiovascular diseases using readily available blood data.
- Identified key hematologic and metabolic indicators for diagnosis and differentiation.
- This cost-effective approach can aid in early diagnosis and prevention efforts.
Background:
Cardiovascular disease, also known as circulation system disease, remains the leading cause of morbidity and mortality worldwide. Traditional methods for diagnosing cardiovascular disease are often expensive and time-consuming. So the purpose of this study is to construct machine learning models for the diagnosis of cardiovascular diseases using easily accessible blood routine and biochemical detection data and explore the unique hematologic features of cardiovascular diseases, including some metabolic indicators.
Methods:
After the data preprocessing, 25,794 healthy people and 32,822 circulation system disease patients with the blood routine and biochemical detection data were utilized for our study. We selected logistic regression, random forest, support vector machine, eXtreme Gradient Boosting (XGBoost), and deep neural network to construct models. Finally, the SHAP algorithm was used to interpret models.
Results:
The circulation system disease prediction model constructed by XGBoost possessed the best performance (AUC: 0.9921 (0.9911-0.9930); Acc: 0.9618 (0.9588-0.9645); Sn: 0.9690 (0.9655-0.9723); Sp: 0.9526 (0.9477-0.9572); PPV: 0.9631 (0.9592-0.9668); NPV: 0.9600 (0.9556-0.9644); MCC: 0.9224 (0.9165-0.9279); F1 score: 0.9661 (0.9634-0.9686)). Most models of distinguishing various circulation system diseases also had good performance, the model performance of distinguishing dilated cardiomyopathy from other circulation system diseases was the best (AUC: 0.9267 (0.8663-0.9752)). The model interpretation by the SHAP algorithm indicated features from biochemical detection made major contributions to predicting circulation system disease, such as potassium (K), total protein (TP), albumin (ALB), and indirect bilirubin (NBIL). But for models of distinguishing various circulation system diseases, we found that red blood cell count (RBC), K, direct bilirubin (DBIL), and glucose (GLU) were the top 4 features subdividing various circulation system diseases.
Conclusions:
The present study constructed multiple models using 50 features from the blood routine and biochemical detection data for the diagnosis of various circulation system diseases. At the same time, the unique hematologic features of various circulation system diseases, including some metabolic-related indicators, were also explored. This cost-effective work will benefit more people and help diagnose and prevent circulation system diseases.
Related Concept Videos
Blood Studies for Cardiovascular System I: Cardiac Biomarkers
The essential diagnostic tools for detecting myocardial necrosis and monitoring individuals suspected of having acute coronary syndrome (ACS) include:
Troponins
Troponins, particularly cardiac troponins I and T, are the most precise and sensitive markers of myocardial injury. They are detectable within 4-6 hours of myocardial injury and remain...
Assessment of the Cardiovascular System I: Subjective Data
Initial Enquiry
Ask the patient about their primary concern and thoroughly explore all reported symptoms.
Medical History
Investigate past illnesses affecting the cardiovascular system, such as angina, anemia, rheumatic fever, congenital heart disease, stroke, thrombophlebitis, dysrhythmias, varicosities
Inquire about symptoms...
Errors occurring during blood pressure monitoring
Several factors...
Blood Studies for Cardiovascular System II: CRP, Hcy, and Cardiac Natriuretic Peptide Markers
These markers indicate stress or strain on the heart muscle:
Natriuretic Peptides (BNP)
Cardiac myocytes produce these hormones in response to ventricular stretching...
Dysrhythmias V: Evaluating Dysrhythmias
Acute Coronary Syndrome III: Diagnostic studies

