Related Experiment Video
Updated: Aug 26, 2026

High-definition Transcranial Direct Current Stimulation over Right Dorsolateral Prefrontal Cortex to Enhance Metacognitive Sensitivity
Published on: September 26, 2025
Neural Value Alignment: Human-AI Collaboration Under Goal-Action Ambiguity
Abstract:
Value alignment plays a crucial role in human-artificial intelligence (AI) collaboration. Traditional approaches attempt to infer human goals from actions to guide AI policies. However, this behavior-level alignment faces an inherent challenge: the ambiguous mapping between goals and actions, as a single action might serve multiple possible goals, while different actions could achieve the same goal. To overcome these limitations, we propose neural value alignment (NVA), a unifying perspective that leverages key variables in human reinforcement learning (RL): reward prediction error (RPE) and state prediction error (SPE). RPE captures outcome discrepancies, refining AI goal inference, while SPE reflects state transition misalignment, shaping AI actions. Using a novel task paradigm that dissociated RPE and SPE, combined with electroencephalography (EEG) recordings, we demonstrated cortical decodability of RPE, SPE, and their co-occurrence, robustly across contexts. Simulations showed that RPE-SPE synergy accelerated value alignment, even under imperfect decoding. This study bridges RL and human-AI interaction, showing that RPE-SPE synergy enables flexible and human-compatible artificial systems operating under goal-action ambiguity.
Related Concept Videos
Impression Management Techniques III: Aligning Actions
The Anchoring-and-Adjustment Heuristic
Automatic Processing and Automatic Social Behavior
Neural Regulation
Nonconscious Mimicry
Decision Making
Automatic decision-making is fast, intuitive, and relies on gut feelings...