Related Experiment Video
Updated: Jun 15, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Feasibility of Artificial Intelligence Powered Adverse Event Analysis: Using a Large Language Model to Analyze
Blair E Warren1,2, Fahd Alkhalifah1,2, Aida Ahrari1,2
1Department of Medical Imaging, University of Toronto, Temerty Faculty of Medicine, Toronto, ON, Canada.
Large language models (LLMs) like GPT-4 can accurately classify interventional radiology microwave ablation device safety events and summarize data, aiding clinicians in analyzing safety information.
Area of Science:
- Medical device safety
- Artificial intelligence in healthcare
- Interventional Radiology
Background:
- Analyzing safety event data for medical devices is crucial for patient safety.
- Interventional radiology (IR) procedures utilize various devices, including microwave ablation systems.
- Manual review of extensive safety data can be time-consuming and resource-intensive.
Purpose of the Study:
- To evaluate the capability of a large language model (LLM), specifically GPT-4, in processing and analyzing interventional radiology microwave ablation device safety event data.
- To determine if an LLM can accurately label, consolidate, and summarize this data, mimicking human analysis.
Main Methods:
- Collected microwave ablation safety data from January 1, 2011, to October 31, 2023.
- Utilized GPT-4 with iterative prompt development for data classification and summarization.
- Employed few-shot learning for accurate event data labeling by malfunction type.
Main Results:
- GPT-4 achieved high accuracy in classifying microwave ablation device data: 96.0% (training), 86.4% (validation), and 87.3% (test).
- Iterative summarization using GPT-4 distilled over 650 reports into a clinically relevant summary, comparable to human interpretation.
- The LLM accurately labeled event data by malfunction type but showed inaccuracies in event class counts.
Conclusions:
- LLMs show feasibility for processing large volumes of IR safety data, serving as a valuable tool for clinicians.
- GPT-4 demonstrated proficiency in accurately labeling device malfunction types through few-shot learning.
- Content distillation via LLMs can generate insightful summaries from extensive safety reports, mirroring human analytical capabilities.
More Related Videos
04:58Reduced Procedure Time and Variability with Active Esophageal Cooling During Radiofrequency Ablation for Atrial Fibrillation
Published on: August 25, 2022
04:09Predicting Treatment Response to Image-Guided Therapies Using Machine Learning: An Example for Trans-Arterial Treatment of Hepatocellular Carcinoma
Published on: October 10, 2018