Related Experiment Video
Updated: May 2, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Performance of a large language model on the reasoning tasks of a physician
Peter G Brodeur1, Thomas A Buckley2, Zahir Kanjee1
1Department of Internal Medicine, Beth Israel Deaconess Medical Center, Boston, MA, USA.
Abstract:
More than 65 years ago, complex clinical diagnostic reasoning cases were introduced as the gold standard for the evaluation of expert medical computing systems, a standard that has held ever since. In this study, we report the results of a physician evaluation of a large language model (LLM) on challenging clinical cases across five experiments with a baseline of hundreds of physicians. We then report a real-world study comparing human expert and artificial intelligence (AI) second opinions in randomly selected patients in the emergency room of a major tertiary academic medical center. In all experiments, the LLM outperformed physician baselines and displayed continued improvement from prior generations of AI clinical decision support. Our study suggests that LLMs have eclipsed most benchmarks of clinical reasoning, motivating the urgent need for prospective trials.
Related Concept Videos
Critical Thinking II
Deductive Reasoning
For example, a researcher can deduce specific predictions...
Patient-centered Care
Inductive Reasoning
Inductive reasoning is common in descriptive science. A life scientist makes observations and records them. This data can be qualitative or...
Language and Cognition
