Related Experiment Video
Updated: Feb 27, 2026

Signal Acquisition, Score Interpretation, and Economics of a Non-Invasive Point-of-Care Test for Coronary Artery Disease
Published on: August 9, 2024
Automatic prediction of coronary artery disease from clinical narratives
Kevin Buchan1, Michele Filannino2, Özlem Uzuner2
1Department of Information Science, State University of New York at Albany, NY, USA.
Insights
This study introduces an automated system to predict coronary artery disease (CAD) using patient medical histories. The novel approach achieves 77.4% F1 score, offering a new method for early CAD detection.
Area of Science:
- Computational medicine and artificial intelligence in healthcare.
- Cardiovascular disease research and predictive analytics.
Background:
- Coronary Artery Disease (CAD) is the leading cause of death globally.
- Clinical free text in medical records contains valuable information for disease prediction.
- Existing research has focused on identifying CAD risk factors, not direct prediction from text.
Purpose of the Study:
- To develop and evaluate an automated system for predicting the development of Coronary Artery Disease (CAD) from clinical free text.
- To establish a novel approach for CAD prediction, marking the first attempt at automatic prediction using narrative medical histories.
- To address the challenge of overfitting in small datasets by employing an ontology-guided feature extraction method.
Main Methods:
- Development of a system to analyze narrative medical histories (clinical free text) for CAD prediction.
- Implementation of an ontology-guided approach for feature extraction to manage a limited feature set and prevent overfitting.
- Comparison of the proposed ontology-guided method with two traditional feature selection techniques on a corpus of diabetic patients.
Main Results:
- The proposed ontology-guided system achieved a state-of-the-art performance.
- The system demonstrated a 77.4% F1 score in predicting coronary artery disease development.
- The ontology-guided feature extraction proved effective in handling small datasets and improving prediction accuracy.
Conclusions:
- Automated prediction of Coronary Artery Disease (CAD) from clinical free text is feasible and effective.
- The ontology-guided feature extraction approach is a promising method for predictive modeling in small medical datasets.
- This system represents a significant advancement in leveraging unstructured medical data for cardiovascular disease prediction.
Abstract:
Coronary Artery Disease (CAD) is not only the most common form of heart disease, but also the leading cause of death in both men and women (Coronary Artery Disease: MedlinePlus, 2015). We present a system that is able to automatically predict whether patients develop coronary artery disease based on their narrative medical histories, i.e., clinical free text. Although the free text in medical records has been used in several studies for identifying risk factors of coronary artery disease, to the best of our knowledge our work marks the first attempt at automatically predicting development of CAD. We tackle this task on a small corpus of diabetic patients. The size of this corpus makes it important to limit the number of features in order to avoid overfitting. We propose an ontology-guided approach to feature extraction, and compare it with two classic feature selection techniques. Our system achieves state-of-the-art performance of 77.4% F1 score.
More Related Videos
Related Concept Videos
Coronary Artery Disease III: Clinical Manifestations
Coronary Artery Disease I: Introduction
Atherosclerosis II: Clinical Manifestations and Diagnostic Tests
Coronary Artery Disease V: Interprofessional Care
Imaging Studies for Cardiovascular System VI: Calcium -Scoring CT
Coronary Artery Disease II: Pathophysiology

