Related Experiment Videos
Automated Identification of Complex Percutaneous Coronary Intervention from Cardiac Catheterization Reports using
Introduction:
Manual abstraction of complex percutaneous coronary intervention (PCI) variables from cardiac catheterization reports is labor-intensive and limits scalable cardiovascular research. We evaluated open-weight large language models (LLMs) for automated complex PCI phenotyping.
Methods:
We evaluated three LLMs (Llama 3.3 70B, Meditron-7B, and BioMistral-7B) using manually annotated catheterization reports from three hospitals within Yale New Haven Health. Models identified PCI reports and extracted six complex PCI features: 3 vessels treated, ≥3 lesions treated, bifurcation PCI with two stents, chronic total occlusion, ≥3 stents, and total stent length ≥60 mm.
Results:
Among 1,412 clinical notes, 596 were PCI reports. Llama 3.3 70B outperformed the smaller domain-specific models across most tasks. For PCI identification, Llama 3.3 70B achieved 100.0% sensitivity, 93.8% specificity, 96.4% accuracy, and 95.9% F1 score. Among 590 evaluable PCI reports (excluding 6 indeterminable cases due to missing variables) for complex PCI classification, Llama 3.3 70B achieved 97.7% sensitivity, 80.1% specificity, 57.6% positive predictive value, 99.2% negative predictive value, 83.9% accuracy, and 72.5% F1 score. Performance was higher for explicitly documented variables, including stent number and length, and lower for variables requiring interpretation across procedural details, including lesion count, bifurcation PCI, and chronic total occlusion PCI. Llama 3.3 70B had the highest accuracy at each site for complex PCI classification but significant site-level heterogeneity was observed.
Conclusion:
A high-capacity open-weight LLM accurately extracted complex PCI variables from unstructured reports and outperformed smaller domain-specific models. These findings support the potential use of locally deployable LLMs for scalable automated PCI phenotyping.