Related Experiment Video
Updated: Jan 13, 2026

Catheter Ablation in Combination With Left Atrial Appendage Closure for Atrial Fibrillation
Published on: February 26, 2013
Leveraging electronic health records for atrial fibrillation cohort generation
Ane G Domingo-Aldama1, Marcos Merino Prado1, Alain García-Olea2
1University of the Basque Country UPV/EHU, 48013 Bilbao, Vizcaya Spain.
Purpose:
Cohort selection and eligibility screening are critical in clinical research, especially in trials where manual patient matching remains a major bottleneck. This study investigates the use of Natural Language Processing and Large Language Models (LLMs) in two real use cases, namely Atrial Fibrillation (AF) progression and Hearth Failure (HF) decompensation, within a non-English clinical context. We specifically address the following research questions: (1) Can discharge reports and NLP support cohort selection? (2) Can LLMs effectively model longitudinal patient trajectories and temporal reasoning? (3) Do general-purpose or domain-adapted LLMs outperform rule-based baselines for this task? (4) Compared to large foundation models, do small-scale LLMs offer similar performance?
Methods:
A dataset of 212 patients was manually annotated for AF progression using discharge reports. Two strategies were evaluated: (1) an adapted rule-based pipeline and (2) zero-shot open-source LLMs with varying prompt structures. To assess generalizability, an additional dataset of 100 patients was annotated for HF decompensation.
Results:
The adapted rule-based approach achieved the highest accuracy (0.82), but LLMs with task-division prompts performed comparably (up to 0.79), requiring significantly less manual effort. The medium-sized general-domain gemma-3 model outperformed others.
Conclusions:
(1) Discharge reports are a valuable resource for automatic cohort selection, with both the rule-based method and LLMs showing promising results. (2) While LLMs struggled with long-context inputs, they handled temporal reasoning well when explicit dates were provided. (3) Larger models did not always outperform smaller ones, (4) prompt language strongly influenced performance, and medical model variants were not consistently superior.
More Related Videos
04:58Reduced Procedure Time and Variability with Active Esophageal Cooling During Radiofrequency Ablation for Atrial Fibrillation
Published on: August 25, 2022
05:03Patient Directed Recording of a Bipolar Three-Lead Electrocardiogram using a Smartwatch with ECG Function
Published on: December 11, 2019
Related Concept Videos
Methods of Documentation VII: EMR
Purpose of Health Records II
Dysrhythmias V: Evaluating Dysrhythmias