Related Experiment Video
Updated: Jul 16, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Retrieval-Augmented Large Language Model for Angiographic Prediction of Coronary Physiology
Sant Kumar1,2, Keshav Nandakumar1, Pedro A Villablanca3
1School of Medicine, Creighton University, Phoenix, AZ 85012, USA.
Abstract:
Background: Invasive physiologic assessment is often needed for moderately stenotic coronary lesions, although it remains underused in routine practice. It remains unknown whether GPT-based large language models (LLMs) can estimate coronary physiology from coronary angiographic images, and whether retrieval augmentation can improve this task. Methods: We performed a retrospective pilot study of consecutive cases undergoing coronary angiography with invasive instantaneous wave-free ratio (iFR) assessment between 2023 and 2025. Eligible cases required invasive iFR and two orthogonal end-diastolic still frames of the target vessel at maximal opacification. We compared a baseline GPT-5.2 model without retrieval-augmented generation (RAG), termed No-RAG, with the same GPT-5.2 model using RAG, termed RAG. Both conditions received identical angiographic frames and structured clinical text. The RAG modification added the top five case-specific text chunks retrieved from four coronary physiology/revascularization documents to provide physiologic thresholds, guideline context, and uncertainty framing; no additional angiographic images or lesion-specific iFR information were provided. Frame-level predictions were averaged to derive a case-level predicted iFR. The primary endpoint was agreement between predicted and measured iFR. Results: Of 34 eligible cases screened, 32 vessels were included. The cohort comprised 25/32 cases (78.1%) with significant disease (iFR ≤ 0.89) and 7 cases classified as non-ischemic. Without RAG, mean absolute error (MAE) was 0.064 and the root mean square error (RMSE) was 0.083, with weak correlation with invasive iFR (r = 0.205, p = 0.259). With RAG, point estimates favored improved continuous agreement, with MAE decreasing to 0.029, RMSE to 0.038, and correlation increasing to r = 0.830 (p < 0.001). Threshold-based classification also yielded higher point estimates for accuracy, increasing from 0.750 to 0.906. Conclusions: In this small pilot study, improved point estimates for agreement between LLM-predicted and invasively measured iFR were seen after adding RAG to a GPT-based model for estimating iFR from angiographic imaging. These findings suggest that functionally classifying coronary stenoses is limited by overestimating severity in less severe stenoses, and that a scaling correction is needed. The results, however, require validation in larger, more balanced cohorts.