Related Experiment Video
Updated: Oct 26, 2025

Investigating the Pathogenesis of MYH7 Mutation Gly823Glu in Familial Hypertrophic Cardiomyopathy using a Mouse Model
Published on: August 8, 2022
Explanatory Analysis of a Machine Learning Model to Identify Hypertrophic Cardiomyopathy Patients from EHR Using
Nasibeh Zanjirani Farahani1, Shivaram Poigai Arunachalam2, Divaakar Siva Baala Sundaram3
1Department of Health Sciences Research, Mayo Clinic, Rochester, MN, USA.
Insights
Machine learning models using electronic health records can improve the diagnosis of hypertrophic cardiomyopathy (HCM), a leading cause of sudden cardiac death. This study developed a predictive model to aid physicians in identifying HCM patients more effectively.
Area of Science:
- Cardiology
- Medical Informatics
- Machine Learning
Background:
- Hypertrophic cardiomyopathy (HCM) is a genetic heart disease and a primary cause of sudden cardiac death in young adults.
- Despite established risk factors and guidelines, HCM diagnosis and management remain suboptimal, leading to underdiagnosis.
- Electronic health record (EHR) data offers potential for developing machine learning models to improve HCM diagnosis.
Purpose of the Study:
- To develop and evaluate a novel predictive model using EHR billing codes for identifying patients with hypertrophic cardiomyopathy (HCM).
- To assist physicians in diagnostic decision-making for HCM through data-driven insights.
- To explore the utility of automated phenotyping using billing codes for HCM detection.
Main Methods:
- A cohort of 11,562 patients with suspected or confirmed HCM from 1995-2019 was analyzed.
- Billing codes were extracted from EHR data, and ground truth labels were established using echocardiography or cardiac magnetic resonance imaging.
- A random forest model was employed to predict HCM status ('definite HCM', 'possible HCM', 'no HCM phenotype').
Main Results:
- The random forest model achieved an accuracy of 71%, weighted recall of 70%, precision of 75%, and weighted F1 score of 72% in identifying HCM patients.
- The model demonstrated effectiveness in classifying patients across 'definite HCM', 'possible HCM', and 'no HCM phenotype' categories.
- Multidimensional scaling and principal component analysis visualizations were generated to aid clinician interpretation.
Conclusions:
- Billing codes within EHR data can be effectively utilized by machine learning models to identify patients with hypertrophic cardiomyopathy (HCM).
- The developed predictive model shows promise in supporting clinical diagnosis and improving patient management for HCM.
- This approach offers a valuable tool for enhancing the identification of HCM through automated EHR phenotyping.
Abstract:
Hypertrophic cardiomyopathy (HCM) is a genetic heart disease that is the leading cause of sudden cardiac death (SCD) in young adults. Despite the well-known risk factors and existing clinical practice guidelines, HCM patients are underdiagnosed and sub-optimally managed. Developing machine learning models on electronic health record (EHR) data can help in better diagnosis of HCM and thus improve hundreds of patient lives. Automated phenotyping using HCM billing codes has received limited attention in the literature with a small number of prior publications. In this paper, we propose a novel predictive model that helps physicians in making diagnostic decisions, by means of information learned from historical data of similar patients. We assembled a cohort of 11,562 patients with known or suspected HCM who have visited Mayo Clinic between the years 1995 to 2019. All existing billing codes of these patients were extracted from the EHR data warehouse. Target ground truth labeling for training the machine learning model was provided by confirmed HCM diagnosis using the gold standard imaging tests for HCM diagnosis echocardiography (echo), or cardiac magnetic resonance (CMR) imaging. As the result, patients were labeled into three categories of "yes definite HCM", "no HCM phenotype", and "possible HCM" after a manual review of medical records and imaging tests. In this study, a random forest was adopted to investigate the predictive performance of billing codes for the identification of HCM patients due to its practical application and expected accuracy in a wide range of use cases. Our model performed well in finding patients with "yes definite", "possible" and "no" HCM with an accuracy of 71%, weighted recall of 70%, the precision of 75%, and weighted F1 score of 72%. Furthermore, we provided visualizations based on multidimensional scaling and the principal component analysis to provide insights for clinicians' interpretation. This model can be used for the identification of HCM patients using their EHR data, and help clinicians in their diagnosis decision making.
More Related Videos
06:55Inverse Probability of Treatment Weighting Propensity Score using the Military Health System Data Repository and National Death Index
Published on: January 8, 2020
05:16Cutoff Value of Phase Angle by Bioelectrical Impedance Analysis at Admission as a Prognostic Factor in Patients with Acute Heart Failure
Published on: June 10, 2025
Related Concept Videos
Cardiomyopathy III: Hypertrophic Cardiomyopathy
Cardiomyopathy V: Interprofessional Care
Cardiomyopathy I: Introduction and Classification
Hypertension III: Clinical Manifestations and Diagnostic Studies
Cardiomyopathy II: Dilated Cardiomyopathy