Related Experiment Video
Updated: Jun 17, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Clinical artificial intelligence: teaching a large language model to generate recommendations that align with
Bright Huo1, Nana Marfo2, Patricia Sylla3
1Division of General Surgery, Department of Surgery, McMaster University, Hamilton, ON, Canada.
Background:
Large Language Models (LLMs) provide clinical guidance with inconsistent accuracy due to limitations with their training dataset. LLMs are "teachable" through customization. We compared the ability of the generic ChatGPT-4 model and a customized version of ChatGPT-4 to provide recommendations for the surgical management of gastroesophageal reflux disease (GERD) to both surgeons and patients.
Methods:
Sixty patient cases were developed using eligibility criteria from the Society of American Gastrointestinal and Endoscopic Surgeons (SAGES) & United European Gastroenterology (UEG)-European Association of Endoscopic. Surgery (EAES) guidelines for the surgical management of GERD. Standardized prompts were engineered for physicians as the end-user, with separate layperson prompts for patients. A customized GPT was developed to generate recommendations based on guidelines, called the GERD Tool for Surgery (GTS). Both the GTS and generic ChatGPT-4 were queried July 21st, 2024. Model performance was evaluated by comparing responses to SAGES & UEG-EAES guideline recommendations. Outcome data was presented using descriptive statistics including counts and percentages.
Results:
The GTS provided accurate recommendations for the surgical management of GERD for 60/60 (100.0%) surgeon inquiries and 40/40 (100.0%) patient inquiries based on guideline recommendations. The Generic ChatGPT-4 model generated accurate guidance for 40/60 (66.7%) surgeon inquiries and 19/40 (47.5%) patient inquiries. The GTS produced recommendations based on the 2021 SAGES & UEG-EAES guidelines on the surgical management of GERD, while the generic ChatGPT-4 model generated guidance without citing evidence to support its recommendations.
Conclusion:
ChatGPT-4 can be customized to overcome limitations with its training dataset to provide recommendations for the surgical management of GERD with reliable accuracy and consistency. The training of LLM models can be used to help integrate this efficient technology into the creation of robust and accurate information for both surgeons and patients. Prospective data is needed to assess its effectiveness in a pragmatic clinical environment.
Related Concept Videos
Gastroesophageal Reflux Disease II: Clinical Features and Management
Clinical Manifestations
GERD presents itself in a multitude of ways, with symptoms varying from person to person. The hallmark symptoms are...
Barrett Esophagus-II: Clinical Manifestations and Management
To diagnose Barrett's esophagus, healthcare providers often recommend an endoscopy for those showing symptoms of acid reflux. The procedure...
Peptic Ulcer Disease V: Surgical Management and Nursing Care
Surgical Interventions for Peptic Ulcer Disease
Esophageal Strictures-II: Clinical Features and Management
Healthcare providers should gather a comprehensive medical history and conduct a physical examination for diagnosis. If esophageal stricture is...
Peptic Ulcer Disease IV: Management
The therapeutic approach involves ensuring adequate rest, implementing drug therapy, promoting smoking cessation, making dietary modifications, and emphasizing long-term follow-up care.
Pharmacological management
The prevailing therapy for peptic ulcers involves a combination of managing the patient's current...
Endoscopic Procedures I: Esophagogastroduodenoscopy
During an EGD, the endoscope can be used to:

