Related Experiment Video
Updated: Jun 11, 2025

Hydra, a Computer-Based Platform for Aiding Clinicians in Cardiovascular Analysis and Diagnosis
Published on: September 26, 2018
Evaluation of Care Quality for Atrial Fibrillation Across Non-Interoperable Electronic Health Record Data using a
Philip Adejumo1,2, Phyllis Thangaraj1,2, Sumukh Vasisht Shankar1,2
1Section of Cardiovascular Medicine, Department of Internal Medicine, Yale School of Medicine, New Haven, CT.
Insights
A novel Retrieval-Augmented Generation (RAG) model accurately extracts stroke risk factors from clinical notes for atrial fibrillation (AF) patients. This enhances risk assessment and guides anticoagulation therapy decisions.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Healthcare
- Clinical Decision Support
Background:
- Accurate stroke risk assessment in atrial fibrillation (AF) is vital for effective anticoagulation therapy.
- Current methods rely on manual CHA₂DS₂-VASc score calculation or limited structured Electronic Health Record (EHR) data.
- Unstructured clinical notes offer rich, underutilized data for improving risk stratification.
Purpose of the Study:
- To develop and validate a Retrieval-Augmented Generation (RAG) approach for extracting CHA₂DS₂-VASc risk factors from unstructured clinical notes in AF patients.
- To enhance the accuracy of stroke risk assessment by leveraging information from free-text clinical narratives.
- To facilitate computable risk assessment for guiding anticoagulation therapy.
Main Methods:
- A RAG architecture, utilizing the Llama3.1 large language model, was employed to extract CHA₂DS₂-VASc risk factors from 1,000 clinical notes.
- A subset of 200 notes was manually annotated by two clinicians to establish a gold standard for validation.
- Performance was assessed using macro-averaged area under the receiver operating characteristic (AUROC), with external validation on MIMIC-IV data.
Main Results:
- The RAG model significantly outperformed structured data in identifying key risk factors like hypertension, stroke/TIA, vascular disease, and diabetes.
- High AUROCs (0.96-0.98) were achieved for hypertension, diabetes, and age ≥75 years in expert-annotated notes.
- Incorporating RAG-identified factors led to increased CHA₂DS₂-VASc scores compared to using structured data alone.
Conclusions:
- Large language model-optimized RAG accurately extracts crucial CHA₂DS₂-VASc risk factors from unstructured AF patient notes.
- This automated approach enables more precise, computable risk assessment.
- The findings support improved guidance for appropriate anticoagulation therapy in AF patients.
Importance:
Standardized assessment of clinical quality measures from electronic health records (EHRs) is challenging because information is fragmented across structured and unstructured data, and due to low interoperability across systems. Traditionally, extracting this information requires manual EHR abstraction, a time-consuming and expensive process that also limits real-time care quality improvement.
Objective:
To evaluate whether a data format-agnostic retrieval-augmented generation-enabled large language model (RAG-LLM) can accurately abstract clinical variables from heterogeneous structured and unstructured EHR data.
Design Setting And Participants:
Retrospective cross-sectional study assessing stroke and bleeding risk in patients with atrial fibrillation (AF) from two health systems. We developed a RAG-LLM model to extract CHA DS -VASc and HAS-BLED risk factors from tabular data and clinical documentation. The framework was validated on 300 expert-annotated patient records (200 from Yale New Haven Health System [YNHHS] and 100 from the Medical Information Mart for Intensive Care [MIMIC-IV]). The system was deployed on two large cohorts: 104,204 patients with AF from YNHHS (2013-2024) and 13,117 from MIMIC-IV (2008-2022). We compared anticoagulation recommendations derived from RAG-LLM with those based on traditional structured data abstraction.
Exposures:
Use of a RAG-LLM model to abstract stroke and bleeding risk factors from structured and unstructured EHR data.
Main Outcomes And Measures:
Accuracy of RAG-LLM-based risk factor abstraction against expert annotation. Secondary outcomes included efficiency, cross-cohort generalizability, and impact on anticoagulation eligibility based on risk stratification.
Results:
In the validation cohort (mean age 74.8 years, 42.7% female), RAG-LLM demonstrated superior performance across all metrics compared with structural data abstraction. For individual CHA DS -VASc components, accuracy ranged from 0.94-1.00 (YNHHS) and 0.89-1.00 (MIMIC-IV) versus 0.66-0.92 (YNHHS) and 0.44-0.97 (MIMIC-IV) for structured data, which was similar for HAS-BLED (0.94-1.00 and 0.89-1.00 vs 0.66-0.94 and 0.44-0.97). In the deployment study, among 3,207 patients classified as low/intermediate stroke risk with structured data, 62.1% (1,993) were reclassified as high risk with RAG-LLM and would become eligible for anticoagulation. Similarly, 5.5% of those classified as low bleeding risk by structured data were reclassified as high risk, substantially refining contraindication assessment.
Conclusions:
A multimodal RAG-LLM accurately abstracts clinical variables from structured and unstructured EHR data to improve stroke and bleeding risk assessments in patients with AF, enhancing identification of appropriate anticoagulation candidates.

