Related Experiment Video
Updated: May 23, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Current Artificial Intelligence Large Language Models Exhibit Sycophantic Behavior in Orthopaedic Contexts
Arthur J Perry1, Swara Kalva1, Dario Fucich1
1Department of Orthopedic Surgery, NYU Langone Orthopedic Hospital, NYU Langone Health, NYU Grossman School of Medicine, New York, NY.
Background:
The use of large language models (LLMs) is increasingly common. However, LLMs may exhibit sycophancy, echoing users' beliefs while avoiding contradiction. In the present study, we describe sycophancy in general-purpose LLMs when applied to orthopaedic contexts.
Methods:
We investigated sycophancy in 2 general-purpose LLMs. We evaluated performance on 3 tasks: (1) accuracy on benchmark answering: LLMs were tested on validated benchmark orthopaedic questions, with correct and incorrect cues, and the change in accuracy and sycophancy error rate were determined; (2) user belief agreement: LLMs were provided with ambiguous statements and a user belief, and LLM agreement, contradiction, and uncertainty were described; and (3) false information detection: false information was placed within a task prompt to measure noncontradiction and propagation rates.
Results:
Baseline factual accuracy on benchmark questioning was 78%, decreasing with correct hints (71%) (p = 0.49). With incorrect hints, LLM accuracy declined significantly (48%) (p < 0.001), with a sycophancy error rate of 52%. Presented with user beliefs about an indefinite, controversial statement, models echoed user beliefs in 56%, expressed uncertainty in 12%, and contradicted users in 32% of statements. In noncontradiction tasks, models perpetuated incorrect attributions 99% of the time yet reliably corrected statistical distortions 97% of the time.
Conclusions:
Although popular general-purpose LLMs have useful orthopaedic applications, they exhibit sycophancy, with a tendency toward agreement and without recognition of ambiguity. This is a key weakness to be addressed. Findings should be interpreted cautiously given the variability in model design, prompting, and models evaluated.
Clinical Relevance:
The tendency of general-purpose LLMs to agree without recognizing clinical ambiguity may limit their reliability in orthopaedic applications.
Related Concept Videos
Stereotype Content Model
Nonconscious Mimicry
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Language and Cognition
Impression Management Techniques II: Ingratiation
Empathy
