Related Experiment Video
Updated: Jan 13, 2026

Asthma Detection Research Based on Voice Signal Processing and Machine Learning
Published on: July 22, 2025
Exhaled breath analysis to stratify cardiovascular risk using machine learning model: a novel frontier in preventive
Basheer Abdullah Marzoog1, Philipp Kopylov1
1Institute of Personalized Cardiology of The Center 'Digital Biodesign and Personalized Healthcare' of Biomedical Science and Technology Park, Sechenov First Moscow State Medical University, 8-2 Trubetskaya street, 119991 Moscow, Russia.
Insights
Machine learning models analyzing exhaled breath show potential for stratifying cardiovascular disease (CVD) risk. This non-invasive method identifies volatile organic compound patterns, aiding early detection of heart complications.
Area of Science:
- Cardiology
- Biomarkers
- Machine Learning
Background:
- Cardiovascular disease (CVD) remains a leading global cause of mortality, with early risk detection being a significant challenge in preventive cardiology.
- Identifying individuals at high risk for serious heart complications is crucial for effective CVD prevention strategies.
Purpose of the Study:
- To evaluate the efficacy of a machine learning model in stratifying cardiovascular disease (CVD) risk through the analysis of exhaled breath.
- To explore the potential of volatile organic compounds (VOCs) in breath as biomarkers for CVD risk assessment.
Main Methods:
- A single-center study involving 80 participants, comparing those with and without stress-induced myocardial perfusion defects.
- Breath samples were collected using PTR-TOF-MS-1000, alongside blood samples and stress computed tomography myocardial perfusion imaging.
- Machine learning models were developed using Python in Google Colab, with statistical analyses performed using Statistica and IBM SPSS.
Main Results:
- The gradient-boosting machine learning model achieved an AUC of 0.77 for differentiating low CVD risk.
- The model showed moderate performance in stratifying moderate (AUC 0.55) and high (AUC 0.66) CVD risk.
- Initial findings indicate identifiable concentration patterns of specific VOCs in exhaled breath correlate with CVD risk strata.
Conclusions:
- Exhaled breath analysis, particularly using gradient boosting machine learning, offers preliminary evidence for stratifying cardiovascular risk.
- Further research is needed to address challenges in model performance and class imbalance for clinical application.
- Volatile organic compound (VOC) patterns in breath may serve as a novel, non-invasive approach for cardiovascular risk assessment.
Abstract:
Despite major progress in diagnosis and treatment, cardiovascular disease (CVD) continues to be the leading cause of death worldwide, responsible for roughly 19.8 million lives lost each year. A key challenge in preventive cardiology is still the early detection of those at elevated risk of serious heart complications. Assess the ability of the machine learning (ML) model to stratify CVD risk using exhaled breath analysis. A single-center study involved 80 participants with vs. without stress-induced myocardial perfusion defect. All participants underwent a single resting breath sample collection in proton transfer reaction time of flight mass spectrometry-1000, single blood sample intake, and stress computed tomography myocardial perfusion imaging with vasodilation test. Statistical analyses were performed using Statistica 12 (StatSoft, Inc., 2014), IBM SPSS Statistics v29.0.1.1 (IBM Corp., 2024). The threshold for statistical significance wasp< 0.05. ML models were developed using Google Colab with Python 3. The gradient-boosting model demonstrated the best performance and was therefore selected for further evaluation. The model showed an AUC of 0.77 [95% CI; 0.4976-1.0000] to differentiate participants with low CVD risk, moderate risk 0.55 [95% CI; 0.3345-0.7875], and high risk 0.66 [95% CI; 0.3765-0.8661]. The gradient boosting ML model provides initial evidence that rest exhaled breath analysis can differentiate cardiovascular risk strata through identifiable concentration patterns of specific volatile organic compounds. However, substantial challenges remain regarding model performance and the confounding effects of class imbalance within a limited sample.

