Related Experiment Video
Updated: Oct 30, 2025

Echocardiographic Assessment Using Subxiphoid-Only Examination for Hypotensive Patients
Published on: April 18, 2025
Supervised Analysis for Phenotype Identification: The Case of Heart Failure Ejection Fraction Class.
Cristina Lopez1, Jose Luis Holgado1, Raquel Cortes1
1Cardiovascular and Renal Research Group, INCLIVA Research Institute, University of Valencia, 46010 Valencia, Spain.
This study developed a machine learning model to identify heart failure patients with reduced heart pumping capacity using electronic health record data. By applying advanced statistical techniques to balance patient groups and select key health indicators, the researchers created an algorithm that accurately flags individuals who may require specialized cardiac care.
Area of Science:
- Artificial intelligence applications within cardiovascular medicine
- Supervised analysis for clinical phenotyping
Background:
Clinical phenotyping through computational clustering represents a significant frontier in modern healthcare innovation. No prior work had resolved the specific challenge of classifying heart failure patients based on left ventricular ejection fraction using routine digital records. That uncertainty drove the need for robust predictive modeling tools. Prior research has shown that electronic health records contain vast amounts of untapped diagnostic information. This gap motivated the development of automated systems to improve patient stratification. It was already known that traditional manual chart reviews are time-consuming and often incomplete. Researchers have long sought ways to leverage existing data to enhance clinical decision-making. This study addresses the pressing requirement for scalable methods to identify high-risk cardiac populations.
Purpose Of The Study:
The study aimed to develop a predictive model for classifying heart failure patients based on their left ventricular ejection fraction. Researchers sought to leverage existing electronic health record data to improve the identification of specific cardiac phenotypes. This effort addresses the challenge of managing patients who may not receive regular follow-up in specialized cardiology clinics. The authors intended to create an automated algorithm capable of flagging individuals with reduced heart pumping capacity. By utilizing supervised learning techniques, they aimed to enhance the accuracy of patient stratification. The team focused on selecting the most influential clinical variables to ensure the model remains robust and interpretable. This work was motivated by the need for scalable solutions to support clinical decision-making in primary care. The researchers aimed to demonstrate that machine learning can effectively bridge the gap between raw clinical data and actionable patient insights.
Main Methods:
The research team implemented a supervised learning design to analyze electronic health record data from over two thousand subjects. Their review approach involved selecting patients with confirmed heart failure and documented echocardiography results. Investigators employed the Least Absolute Shrinkage and Selection Operator to isolate the most impactful clinical variables. To mitigate data skewness, they utilized the Synthetic Minority Oversampling Technique during the training phase. The study constructed two distinct predictive frameworks, specifically Random Forest and XGBoost models. These computational tools underwent rigorous testing to evaluate their classification performance. The final algorithm was subsequently applied to a larger cohort of over twenty-five thousand primary care patients. This systematic process ensured the development of a reliable tool for identifying specific cardiac phenotypes.
Main Results:
Key findings from the literature indicate that the full XGBoost model achieved the highest predictive performance among the tested architectures. This model demonstrated superior accuracy alongside high positive and negative predictive values. The analysis identified gender, age, unstable angina, atrial fibrillation, and acute myocardial infarct as the most influential predictors of ejection fraction. When applied to the primary care dataset, the algorithm successfully flagged 6,170 individuals, representing 21.1 percent of the total cohort. These patients were identified as belonging to the reduced ejection fraction group. The model effectively processed records from patients lacking regular cardiology clinic follow-up. These results confirm the utility of machine learning in extracting actionable insights from routine clinical documentation. The data suggests a strong correlation between the selected variables and the target cardiac phenotype.
Conclusions:
The researchers propose that their predictive model effectively identifies heart failure patients with reduced pumping capacity. This approach offers a viable pathway for improving patient management in primary care settings. The authors suggest that their methodology remains highly adaptable for future studies utilizing large-scale electronic health record datasets. Their findings demonstrate that specific clinical variables, such as age and atrial fibrillation, exert a strong influence on ejection fraction status. The team notes that the algorithm successfully flags individuals who might otherwise lack regular cardiology follow-up. This work highlights the potential for machine learning to optimize resource allocation within healthcare systems. The authors conclude that their protocol provides a structured way to prioritize patients for specialized interventions. Their results provide a foundation for integrating automated phenotyping into routine clinical workflows.
Frequently Asked Questions
The researchers propose a supervised machine learning approach using Random Forest and XGBoost models. By integrating LASSO variable selection and SMOTE for data balancing, the algorithm achieves high predictive accuracy in identifying patients with reduced ejection fraction.
The team utilized the Synthetic Minority Oversampling Technique to address class imbalance. This method ensures that the model does not become biased toward the majority group, allowing for more reliable detection of the minority class within the electronic health record dataset.
The authors state that echocardiography measurements are necessary to define the gold standard for ejection fraction. This clinical data serves as the target variable for training the algorithm, ensuring that the model predictions align with established cardiac diagnostic criteria.
Electronic health records serve as the primary data source for training and validation. These records provide the necessary longitudinal information, including ICD-codes and clinical history, which allows the model to identify patients who lack regular specialized cardiology follow-up.
The researchers measured model performance using accuracy, negative predictive value, and positive predictive value. These metrics confirm that the XGBoost model outperforms other tested approaches in correctly identifying the target patient population.
The authors imply that this methodology could facilitate broader implementation of automated screening protocols. By identifying patients with reduced ejection fraction, healthcare systems can potentially improve clinical outcomes through targeted, timely interventions for those at higher risk.
Related Concept Videos
Heart Failure IV: Classification and Diagnostic Evaluation
Pathophysiology of Heart Failure
Heart Failure II: Pathophysiology
Cardiomyopathy III: Hypertrophic Cardiomyopathy

