Related Experiment Video
Updated: May 19, 2026

Virtual Agent for Real-Time Motivational Interviewing by Integrating Adaptive Nonverbal Behavior and Language Models
Published on: December 23, 2025
Simple picture of how output from ChatGPT-like AI shifts from good to bad
Neil F Johnson1, Frank Yingjie Huo1
1Physics Department, George Washington University, Washington, DC 20052, USA.
Abstract:
Generative language models can drift mid-generation from reliable, helpful continuations to more undesirable, misleading, or unsafe ones. Such shifts can be hard to notice because the early part of a response can be fluent and correct, so downstream users (or automations) may not trigger checks until after harm occurs. Here, we isolate a minimal, first-principles mechanism for such "good-to-bad" output shifts. The mechanism originates within a single effective attention head, where the dot products of embedded vectors drive a competition between output basins; mapping to a multispin thermal system yields a closed-form expression for the shift. Crucially, multilayer processing does not wash out this single-head mechanism-it amplifies it. Tracking hidden states through all layers of open-weight models, we find that the competing dot-product gaps grow by up to from the first layer to the penultimate layer, and that successive layers fuse competing basins into a shared geometric subspace-constructing precisely the configuration the single-head formula assumes. A cross-architecture study on six independently trained models confirms that the resulting closed-form expression, with zero free parameters, correctly predicts the output-shift regime in of nonambiguous cases. Our physical picture is of course extremely simplified-just as a paper plane cannot capture all the details of modern aviation, or a simple model of an atom cannot capture all the details of a complex material-but it helps fill a current need for generative AI explainability, and it provides a simple yet concrete platform for discussing how training, fine-tuning and prompts might shift generative AI output.
Related Concept Videos
Automatic Processing and Automatic Social Behavior
Positive and Negative Feedback Loops
Effects of feedback
Feedback significantly modifies the gain of a control system. The gain of a system without feedback is altered by a factor of one plus GH, where G represents...
Negative and Positive Feedback