Related Experiment Video
Updated: May 6, 2026

Ultrasound Imaging of the Thoracic and Abdominal Aorta in Mice to Determine Aneurysm Dimensions
Published on: March 8, 2019
Large language models accurately extract aortic information from abdominal imaging reports in a large, real-world
Colleen P Flanagan1, Lawrence D Gerstley2, Steven Okuhn3
1Division of Vascular & Endovascular Surgery, University of California San Francisco, San Francisco, CA; Division of Clinical Informatics & Digital Transformation, University of California San Francisco, San Francisco, CA.
Large language models (LLMs) accurately extract abdominal aortic aneurysm (AAA) data from imaging reports without task-specific training. This AI approach offers a cost-effective solution to reduce administrative burdens in surveillance programs.
Area of Science:
- Radiology and Artificial Intelligence
- Medical Informatics
- Cardiovascular Imaging Analysis
Background:
- Abdominal aortic aneurysm (AAA) surveillance requires extensive manual data review, making it costly and labor-intensive.
- Existing natural language processing (NLP) tools need task-specific human training for each application.
- Generalized artificial intelligence offers a potential solution to automate data extraction without prior training.
Purpose of the Study:
- To evaluate the efficacy of a large language model (LLM) in extracting AAA-related data from abdominal imaging reports.
- To determine if LLM can perform data extraction without task-based human training, unlike traditional NLP methods.
- To assess the reliability and efficiency of LLM for AAA surveillance data mining.
Main Methods:
- Utilized Llama 3.3 70B LLM on a local server to analyze 16,331 abdominal imaging reports (ultrasound, CT, MRI) from an AAA surveillance registry (2008-2024).
- LLM extracted maximal abdominal aortic diameter or determined AAA presence/status based on descriptive terms.
- Compared LLM extracted data against independent expert reviewers using accuracy, sensitivity, positive predictive value, and F1-score metrics.
Main Results:
- The LLM achieved an overall accuracy of 0.93, sensitivity of 0.96, precision of 0.96, and F1-score of 0.96.
- A discrete maximal abdominal aortic diameter was present in 51.9% of reports.
- For AAA diameters between 3-7 cm, the LLM's F1 score ranged from 0.97 to 0.99, demonstrating high reliability.
Conclusions:
- LLM extraction of aortic information from abdominal imaging reports is highly reliable and does not require additional human-directed training.
- LLMs provide a flexible, efficient, and cost-effective method for data mining in AAA surveillance.
- This AI approach can significantly reduce administrative burdens and enhance the quality and efficiency of clinical research in cardiovascular imaging.

