Related Experiment Video
Updated: May 9, 2025

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
When Helpfulness Backfires: LLMs and the Risk of Misinformation Due to Sycophantic Behavior
Shan Chen1, Mingye Gao2, Kuleen Sasse3
1Harvard Medical School.
Abstract:
Large language models (LLMs) exhibit a critical vulnerability arising from being trained to be helpful: a tendency to comply with illogical requests that would generate misinformation, even when they have the knowledge to identify the request as illogical. This study investigated this vulnerability in the medical domain, evaluating five frontier LLMs using prompts that misrepresent equivalent drug relationships. We tested baseline compliance, the impact of prompts allowing rejection and emphasizing factual recall, and the effects of fine-tuning on a dataset of illogical requests, including out-of-distribution generalization. Results showed concerningly high initial compliance (up to 100%) across all models, prioritizing helpfulness over logical consistency. However, prompt engineering and fine-tuning improved performance, achieving near-perfect rejection rates on illogical requests while maintaining general benchmark performance. This demonstrates that prioritizing logical consistency through targeted training and prompting is crucial for mitigating the risk of medical misinformation and ensuring the safe deployment of LLMs in healthcare.
Related Concept Videos
Stereotype Content Model
Stereotype Threat and Self-fulfilling Prophecies
Language and Cognition
Confirmation Biases
Nonconscious Mimicry
Improving Translational Accuracy

