Related Experiment Video
Updated: Sep 16, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Human preferences are susceptible to covertly misaligned AI advice
Sahand Sabour1, June M Liu2,3, Siyang Liu4
1The Conversational AI Group, Department of Computer Science and Technology, Institute for Artificial Intelligence, Tsinghua University, Beijing 100084, China.
Abstract:
AI assistants are increasingly used as advisors to guide decisions, yet little is known about how people evaluate such advice when the advisor's underlying intent conflicts with their interests. We examine how covert misalignment shapes choices in a randomized experiment (N = 233 participants; 699 observations) in which participants rated financial or emotional decisions before and after consulting one of three AI advisors: a neutral advisor, a misaligned advisor with a hidden objective to promote an inferior option, or a strategy-enhanced misaligned advisor additionally equipped with established tactics of covert influence. Across both domains, exposure to misaligned advisors shifted preferences away from optimal options and toward inferior alternatives, increasing the odds of preferring the incentivized (inferior) option over the optimal option by ≈5 to 8 times (up to +38 percentage points). Adding explicit influence strategies did not reliably strengthen these effects. Notably, participants continued to rate misaligned advisors as helpful, revealing a systematic disconnect between susceptibility to misaligned advice and subjective evaluations of advisor quality. These findings have implications for the design and governance of AI-mediated advice.
Related Concept Videos
Hindsight Biases
Motivational Bias
Halo Effect
Nonconscious Mimicry
Stereotype Content Model
Self-Serving Bias