Related Experiment Video
Updated: Jan 10, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Benchmarking proprietary and open-source language and vision-language models for gastroenterology clinical reasoning
Seyed Amir Ahmad Safavi-Naini1,2,3, Shuhaib Ali4, Omer Shahab5
1Division of Data-Driven and Digital Health (D3M), The Charles Bronfman Institute for Personalized Medicine, Icahn School of Medicine at Mount Sinai, New York, NY, USA.
Large language models (LLMs) and vision-language models (VLMs) show promise in gastroenterology. Proprietary models like o1-preview and Claude 3.5 Sonnet outperformed open-source options, while quantized models performed comparably to full-precision ones.
Area of Science:
- Medical Artificial Intelligence
- Gastroenterology Applications
- Natural Language Processing
Background:
- The integration of advanced AI, including large language models (LLMs) and vision-language models (VLMs), is rapidly expanding across medical disciplines.
- Assessing the performance of these models in specialized fields like gastroenterology is crucial for understanding their potential clinical utility.
Purpose of the Study:
- To evaluate the effectiveness and accuracy of various proprietary and open-source LLMs and VLMs in a gastroenterology context.
- To compare model performance across different interfaces, computing environments, and quantization levels.
Main Methods:
- Utilized board-style, multiple-choice questions relevant to gastroenterology to test LLMs and VLMs (e.g., GPT, Claude, Gemini, Mistral, Llama, Mixtral, Phi, Qwen).
- Assessed performance under varying conditions, including different compression (quantization) levels for open-source models.
- Investigated the impact of image captions (human-generated, original, LLM-generated) on VLM performance for image-based questions.
Main Results:
- Proprietary models o1-preview (82.0%) and Claude3.5-Sonnet (74.0%) achieved the highest accuracy.
- Top open-source models Llama3.3-70b (65.7%) and Qwen-2.5-72b (61.0%) were outperformed by proprietary models.
- Small, quantized open-source models (8-bit Llama 3.2-11b, 6-bit Phi3-14b) showed performance comparable to their full-precision versions.
- Vision-language model accuracy increased with human-generated captions but decreased with LLM-generated captions for image-based questions.
Conclusions:
- Proprietary LLMs and VLMs demonstrate superior performance in gastroenterology assessments compared to current open-source models.
- Quantization offers a viable method for reducing the computational requirements of open-source models without significant accuracy loss.
- The quality of image captions critically influences VLM performance, highlighting the need for careful data preparation in clinical applications.
Related Concept Videos
Imaging Studies III: Gastrointestinal Motility Studies and Virtual Colonoscopy
Radionuclide Testing
Radionuclide testing is a sophisticated medical technique for assessing gastrointestinal motility. It focuses on gastric emptying and colonic transit time. Radioactive markers track the movement of food through the digestive system, providing insights into gastrointestinal disorders.
In gastric emptying studies, a meal's liquid and...
Assessment of the Gastrointestinal System II: Health Perception Pattern
Health Perception Patterns
Health perception patterns offer valuable insights into a patient's lifestyle habits and how they may impact their GI health. These patterns include:
Pharmacokinetic Models: Comparison and Selection Criterion
Physiological models take a detailed approach by considering specific molecular processes. They can predict drug distribution, metabolism, and elimination changes, providing a comprehensive understanding of how drugs interact with the body.
Language and Cognition

