Related Experiment Video
Updated: May 23, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Use of large language models as clinical decision support tools for management pancreatic adenocarcinoma using
Kristen N Kaiser1, Alexa J Hughes2, Anthony D Yang3
1Surgical Outcomes and Quality Improvement Center (SOQIC), Department of Surgery, Indiana School of Medicine, Indianapolis, IN. Electronic address: https://twitter.com/kristen_kaiser1.
Background:
Large language models may form the basis of clinical decision support tools to improve rates of guideline concordant care for pancreatic ductal adenocarcinoma. The objectives of this study were to 1) define the first-pass accuracy of 2 publicly available large language models in responding to prompts on the basis of National Comprehensive Cancer Network guidelines for pancreatic ductal adenocarcinoma, 2) describe consistency of responses within each large language models, and 3) explore differences between the 2 large language models in their accuracy and verbosity.
Methods:
Clinical scenarios were developed on the basis of current National Comprehensive Cancer Network guidelines. Scenario prompts were entered independently by 2 investigators into OpenAI ChatGPT and Microsoft Copilot, yielding 4 responses per scenario. Responses were manually graded on accuracy and verbosity and compared to clinician-derived responses.
Results:
From the 104 responses, large language model responses were graded as completely correct in 42% of responses (n = 44). ChatGPT responses were more accurate than Copilot across all prompts (3.33 ± 0.86 vs 3.02 ± 0.87, P = .04). Among 54 generated responses from ChatGPT sessions, 52% (n = 27) were completely correct, 35% (n = 18) contained missing information, and 14% (n = 7) were inaccurate/misleading. Copilot responses were completely correct in 33% (n = 17) of responses, whereas 42% (n = 22) were missing information and 25% (n = 13) contained inaccurate/misleading information. Clinician responses were more concise than all large language model-generated responses (32 ± 13 vs 270 ± 70 words, P < .001).
Conclusion:
Large language model-powered responses to clinical questions regarding pancreatic ductal adenocarcinoma are often inaccurate and verbose. These publicly available large language models require significant optimization before implementation within health care as clinical decision support tools.
Related Concept Videos
Chronic Pancreatitis II: Collaborative Care
Assessment:
Cancer Survival Analysis

