Related Experiment Video
Updated: Jun 4, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models can accurately populate Vascular Quality Initiative procedural databases using narrative
Colleen P Flanagan1, Karen Trang2, Joyce Nacario3
1Division of Vascular and Endovascular Surgery, Department of Surgery, University of California San Francisco, San Francisco, CA; Division of Clinical Informatics and Digital Transformation, Department of Medicine, University of California San Francisco, San Francisco, CA.
Large language models (LLMs) can accurately populate Vascular Quality Initiative (VQI) databases from operative reports, improving surgical data entry efficiency and potentially increasing VQI participation.
Area of Science:
- Vascular surgery data management
- Artificial intelligence in healthcare
- Health informatics
Background:
- Vascular Quality Initiative (VQI) participation offers valuable resources but is often hindered by time and personnel constraints for data entry.
- Large language models (LLMs) demonstrate potential in natural language processing and text generation, offering a solution to data entry challenges.
Purpose of the Study:
- To evaluate the accuracy of LLMs in populating VQI procedural databases using operative reports.
- To assess the feasibility of using generative AI to streamline data entry for vascular surgery quality initiatives.
Main Methods:
- A retrospective study analyzed 150 operative reports for carotid endarterectomy (CEA), endovascular aneurysm repair (EVAR), and infrainguinal lower extremity bypass (LEB) procedures.
- A HIPAA-compliant LLM (Versa, based on ChatGPT) was used to automatically extract and populate VQI data from reports.
- The accuracy of two models, gpt-35-turbo and gpt-4, was compared against existing VQI data, with a metric defined as 'unavailable' if discussed in less than 20% of reports.
Main Results:
- The gpt-35-turbo model achieved median accuracy rates of 84.0% for CEA, 92.2% for EVAR, and 84.3% for LEB.
- Excluding routinely unavailable metrics, accuracy increased to 95.5% for CEA, 94.8% for EVAR, and 93.2% for LEB.
- Gpt-4 did not significantly improve performance over gpt-35-turbo, and processing costs were minimal ($0.12 for gpt-35-turbo vs. $3.39 for gpt-4 for 150 reports).
Conclusions:
- LLMs can accurately populate VQI databases with structured and unstructured data at a low cost.
- Increased workflow efficiency through LLMs may enhance a center's ability to participate in the VQI.
- Further research is warranted to explore other VQI databases and improve LLM accuracy for surgical data management.
Related Concept Videos
Pre-Procedural Guidelines for Assessing Blood Pressure
Types of Reports III: Telephone and Verbal Reports
Here's an overview of each type:
Telephone Orders
Nursing Clinical Information System
A Nursing Clinical Information System (NCIS) is a specialized type of healthcare information system tailored to meet the unique needs of nursing practice. It incorporates the principles of nursing informatics to streamline information management and improve the quality of care delivery.
Critical attributes of NCIS include:
Formats for Nursing Documentation
Nursing Assessment Form:
• A nursing assessment form is a foundational document that captures detailed patient data from physical assessments and nursing histories.
• It includes patient demographics, medical history,...
Methods of Documentation VI: Case Management Model
For example, a patient with a chronic...
Guidelines for Nursing Documentation I
Factual:
The following points emphasize the significance of upholding accurate and unbiased documentation in healthcare.

