Related Experiment Video
Updated: May 21, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
474
Performance of Plug-In Augmented ChatGPT and Its Ability to Quantify Uncertainty: Simulation Study on the German
Julian Madrid1, Philipp Diehl1, Mischa Selig2,3
1Department of Cardiology, Pneumology, Angiology, Acute Geriatrics and Intensive Care, Ortenau Klinikum, Klosterstrasse 18, Lahr, 77933, Germany, 49 7821932403.
JMIR Medical Education
|March 21, 2025
Summary
Large language models like GPT-4 demonstrate medical board competency, exceeding minimum thresholds. While they can quantify response uncertainty, overconfidence remains a challenge for clinical AI applications.
Area of Science:
- Artificial Intelligence in Medicine
- Medical Education Technology
- Natural Language Processing
Background:
- Large language models (LLMs) like GPT-4 show increased interest and proposed applications.
- Current LLMs face limitations in symbolic representation and accessing real-time data.
- GPT-4 with plugins aims to address some of these limitations.
Purpose of the Study:
- To evaluate GPT-3.5, GPT-4, and GPT-4 with plugins on the German medical board examination.
- To assess LLM's ability to quantify uncertainty in medical contexts.
- To develop a 'confidence accuracy' metric for evaluating LLM uncertainty quantification.
Main Methods:
- GPT models answered questions from the German medical board examination.
- Analysis included answer justification, response accuracy, and error structure.
- Bootstrapping and confidence intervals assessed statistical significance.
Main Results:
- GPT models surpassed the minimum competency threshold for medical board certification.
- Models demonstrated uncertainty quantification but exhibited overconfidence.
- Distinct justification and reasoning patterns were observed in GPT-generated answers.
Conclusions:
- High performance of GPTs in medical question answering suggests academic and clinical potential.
- Uncertainty quantification capability positions AI as a valuable clinical decision-making tool.
- Significant challenges remain for robust and safe AI implementation in the medical field.

