Related Experiment Video
Updated: May 29, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Multiple large language models versus experienced physicians in diagnosing challenging cases with gastrointestinal
Xintian Yang1, Tongxin Li1, Han Wang2
1State Key Laboratory of Holistic Integrative Management of Gastrointestinal Cancers and National Clinical Research Center for Digestive Diseases, Xijing Hospital of Digestive Diseases, Fourth Military Medical University, Xi'an, China.
Large language models (LLMs) show promise in assisting doctors with challenging diagnoses. Claude 3.5 Sonnet outperformed human gastroenterologists in diagnosing complex gastrointestinal cases.
Area of Science:
- Medical Diagnostics
- Artificial Intelligence in Healthcare
- Gastroenterology
Background:
- Physicians increasingly consult large language models (LLMs) for complex diagnostic challenges.
- The diagnostic capabilities of LLMs compared to human experts remain an area of active research.
Purpose of the Study:
- To compare the diagnostic accuracy of LLMs against human gastroenterologists for challenging gastrointestinal cases.
- To evaluate the effectiveness of LLMs as a diagnostic aid in clinical practice.
Main Methods:
- An offline dataset of 67 challenging gastrointestinal cases was utilized.
- Seven LLMs and 22 gastroenterologists were tasked with providing diagnoses for the cases.
- Diagnostic coverage and instructive diagnoses were compared between LLMs and human physicians.
Main Results:
- Claude 3.5 Sonnet provided the highest proportion of instructive diagnoses (76.1%), significantly outperforming all participating gastroenterologists.
- Claude 3.5 Sonnet's diagnostic coverage rate (76.1%) was significantly higher than that of gastroenterologists using traditional resources (45.5%).
- The performance difference was statistically significant (p < 0.05 for gastroenterologist comparison, p < 0.001 for resource comparison).
Conclusions:
- Advanced LLMs, such as Claude 3.5 Sonnet, demonstrate superior diagnostic capabilities for challenging gastrointestinal cases compared to human experts.
- LLMs can serve as valuable, time-saving, and cost-effective tools to assist gastroenterologists in complex diagnostic scenarios.
- The findings suggest a potential role for AI in enhancing diagnostic accuracy and efficiency in specialized medical fields.
Related Concept Videos
Imaging Studies III: Gastrointestinal Motility Studies and Virtual Colonoscopy
Radionuclide Testing
Radionuclide testing is a sophisticated medical technique for assessing gastrointestinal motility. It focuses on gastric emptying and colonic transit time. Radioactive markers track the movement of food through the digestive system, providing insights into gastrointestinal disorders.
In gastric emptying studies, a meal's liquid and...
Assessment of the Gastrointestinal System I: Subjective Data
Health History
The initial step in assessing the GI system is obtaining a comprehensive health history. This includes inquiring about the patient's history or presence of problems...
Peptic Ulcer Disease III: Clinical Manifestations and Diagnostic Studies
Few clinical manifestations differentiate gastric ulcers from duodenal ulcers. Distinctions in the location, timing, and pain relief are crucial for healthcare providers in differentiating between gastric and duodenal ulcers during clinical assessments.
Assessment of the Gastrointestinal System II: Health Perception Pattern
Health Perception Patterns
Health perception patterns offer valuable insights into a patient's lifestyle habits and how they may impact their GI health. These patterns include:
Serum Laboratory Studies, Stool Test, Breath Test
Endoscopic Procedures IV: Sigmoidoscopy and Laproscopy
Sigmoidoscopy
Sigmoidoscopy is a diagnostic procedure that uses a flexible sigmoidoscope equipped with a light source and camera to examine the rectum and sigmoid colon. The procedure involves inserting the tube through the anus...

