Related Experiment Videos
Machine Learning for Cardiovascular Prevention Prescriptions: Real-World vs. Synthetic Data
Alaedine Benani1,2,3, Damien Grosgeorge1, Pierre Bauvin1
1Preventive Medicine, Data Science and AI Lab, Zoī.
None:
Cardiovascular diseases remain the leading cause of death worldwide, primarily driven by atherosclerosis, which is targeted by lipid-lowering agents such as statins and berberine. This study investigates the use of machine learning (ML) to predict expert-level prescriptions of these therapies using both real-world clinical data and synthetic datasets generated with two approaches: Avatar-based (SD_avatar) and Statistical Data Vault (SD_sdv). After automatic and manual feature selection, multiple ML and neural network models were trained and evaluated through cross-validation using the F1-score as the main performance metric. The best-performing configurations were subsequently tested on a held-out dataset. Real-world data (RD) achieved the highest predictive performance (F1 = 0.67), comparable to both the hybrid dataset (RD + SD_avatar, F1 = 0.67) and the SD-avatar dataset alone (F1 = 0.67), while SDV-based data performed significantly lower (F1 = 0.39). These findings confirm the feasibility of modeling prescription behavior using ML, underline the potential of avatar-generated synthetic data, and highlight the remaining gap between structured data and the nuanced reasoning underlying expert clinical prescriptions.
Related Concept Videos
Cardiovascular Drugs: Classification based on Therapeutic Indications
Coronary Artery Disease IV: Preventive Measures
Heart Failure Drugs: β-Blockers