Related Experiment Video
Updated: Jun 25, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models exhibit greater diagnostic anchoring than physicians in a forced-choice vignette study
Alexander Sheppert1, Christine Shen1, Erik Geissal1
1Legacy Health, 2211 NE 139th Street, Vancouver, WA 98686, USA; The Vancouver Clinic, 700 NE 87th Avenue, Vancouver, WA 98664, USA.
Large language models (LLMs) show greater diagnostic anchoring bias than human physicians in controlled studies. This highlights the need for careful evaluation before integrating LLMs into clinical decision-making to prevent diagnostic errors.
Area of Science:
- Medical Informatics
- Artificial Intelligence in Medicine
- Clinical Reasoning
Background:
- Anchoring bias is a known factor contributing to diagnostic errors in human clinicians.
- The susceptibility of large language models (LLMs) to diagnostic anchoring bias is not well understood.
- This study investigates LLM performance in diagnostic tasks compared to human physicians.
Purpose of the Study:
- To compare the diagnostic anchoring susceptibility of LLMs and human physicians.
- To quantify the difference in anchoring bias between LLMs and physicians in a controlled setting.
- To inform the safe integration of LLMs into clinical diagnostic workflows.
Main Methods:
- A forced-choice ranking study using internal medicine clinical vignettes was conducted.
- Nine vignette pairs compared anchored (suggestive diagnosis) and control conditions.
- Twenty residents, five attending physicians, and eight LLMs evaluated differential diagnoses.
Main Results:
- LLMs ranked the anchor diagnosis first in 55.6% of cases, significantly higher than residents (21.2%) and attendings (10.0%).
- LLMs demonstrated higher odds of anchoring bias compared to both resident and attending physicians.
- LLMs included anchors in the top-5 diagnoses more frequently (97.2%) than physicians (65.0-67.5%).
Conclusions:
- LLMs exhibit significantly greater diagnostic anchoring bias than human physicians in this vignette-based study.
- Findings suggest a need for rigorous evaluation of LLM anchoring susceptibility before clinical deployment.
- Further research in real-world clinical settings is essential to understand LLM performance and safety.
Related Concept Videos
The Anchoring-and-Adjustment Heuristic
Language and Cognition
The Availability Heuristic
Lateralization