Related Experiment Video
Updated: Jan 9, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Use of Large Language Models to Determine the Surveillance Colonoscopy Interval: A Bi-Institutional Validation Study
Vedant Acharya1, Shivan J Mehta2, Daniel A Sussman3
1Department of Radiology, University of Pennsylvania Perelman School of Medicine, Philadelphia, Pennsylvania, USA.
Introduction:
To determine the appropriate postpolypectomy colonoscopy surveillance interval, endoscopists synthesize information from colonoscopy and pathology report impressions and subsequently apply guideline-recommended interval algorithms, such as those developed by the United States Multi-Society Task Force. Given the complexity of these guidelines, this manual process is error-prone, necessitating automated tools, including large language models (LLMs), to improve guideline adherence. The primary aim of this study was to identify the LLM performance in determining the guideline-concordant postpolypectomy surveillance interval on a cohort of 1,000 real-world colonoscopy and pathology report impressions.
Methods:
The data of patients who underwent a screening or surveillance colonoscopy in 2023-2024 at 2 academic health centers were included. Using a custom prompt outlining the US Multi-Society Task Force postpolypectomy surveillance algorithm, the LLM (GPT-4o) was asked to determine the appropriate surveillance interval for all 1,000 examples in the data set. This experiment, using the same model, prompt, and data set, was repeated 10 times; all experiments were conducted between January 27, 2025, and February 3, 2025.
Results:
Across 10 experiments, the average accuracy was 94.6%. There was no significant difference in accuracy based on the institution from which the data originated or the presence of synchronous upper gastrointestinal endoscopy data within the pathology report impression. Examples with 1-3 colon polyps had an average accuracy of 95.8% whereas examples with 4 or more colon polyps had an average accuracy of 88.2%, combined P value < 0.001.
Discussion:
LLMs with a custom prompt achieve consistently high accuracy in determining the guideline-based surveillance colonoscopy interval.
More Related Videos
07:13Comparison of Predictive Performance of Three Lymph Node Staging Systems in Colorectal Signet Ring Cell Carcinoma Based on Machine Learning Model
Published on: April 18, 2025
15:49Flexible Colonoscopy in Mice to Evaluate the Severity of Colitis and Colorectal Tumors Using a Validated Endoscopic Scoring System
Published on: October 16, 2013
Related Concept Videos
Imaging Studies III: Gastrointestinal Motility Studies and Virtual Colonoscopy
Radionuclide Testing
Radionuclide testing is a sophisticated medical technique for assessing gastrointestinal motility. It focuses on gastric emptying and colonic transit time. Radioactive markers track the movement of food through the digestive system, providing insights into gastrointestinal disorders.
In gastric emptying studies, a meal's liquid and...
Improving Translational Accuracy
Improving Translational Accuracy