Related Experiment Video
Updated: Jun 12, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large language models as experimental systems in human psychopathology: a modelling study
Magdalena Katharina Wekenborg1, Elizabeth Anna Mathilde Michels1, Georg Kurze1
1Else Kroener Fresenius Center for Digital Health, Faculty of Medicine and University Hospital Carl Gustav Carus, TUD Dresden University of Technology, Dresden, Germany.
Background:
Despite advances in biomedical research, human psychopathology remains underserved by experimental model systems, limiting therapeutic innovation. Alternative approaches are needed to investigate the mechanisms underlying mental health conditions. We aimed to assess whether large language models (LLMs) could serve as experimental systems to model affective processes relevant to human psychopathology.
Methods:
Using standard psychological induction protocols, including imagery vignettes, we tested whether seven affective states (fear, anxiety, anger, disgust, sadness, worry, and stress) could be systematically induced in six state-of-the-art LLMs (including GPT-4o and several Llama variants) and subsequently reversed with regulation strategies. Separate prompt sequences were created for each affective state, and an additional neutral control condition was included in which no affective stimulation was applied after the initial general prompt. All affective states followed static vignette protocols without interaction, consistent with human studies. The only exception was stress, which, in line with the original Trier Social Stress Test (TSST) protocol, followed a dynamic interactive prompting procedure. The LLMs were intermittently prompted to self-assess their current affective state via visual analogue scales with a fixed numerical range from 0 to 100 (except for anxiety, which was measured with the State version of the State-Trait Anxiety Inventory). To reverse the induction of affective states, a mindfulness-based relaxation technique was used for all conditions except stress, which was followed by a standardised debriefing procedure consistent with the original TSST protocol. Each condition was repeated in five independent runs to ensure reliability. A sentence completion test was used to test for cognitive bias after induction of sadness, with responses rated for emotional valence by three independent human raters. Inter-rater reliability (Cohen's κ) and negativity scores were calculated, and conditions were compared using t tests (Cohen's d).
Findings:
Across all affective states and over five runs, mean scores for GPT-4o increased by 52·83 points (201·20%) from baseline to post affect induction, and downregulation prompts reduced scores by 48·23 points (60·98%). These patterns were broadly replicated across five additional open-weight LLMs, with significant between-model differences for all affective states (p values between 0·045 and <0·0001) except stress (p=0·063). GPT-4o and Llama 4 Maverick showed the strongest effects, whereas Llama 4 Scout showed the weakest responses, indicating that model architecture and scale influence susceptibility to affect induction. In the test for cognitive bias, sadness-related prompts elicited a consistent negativity bias in sentence completions by GPT-4o compared with neutral prompts (mean 15·00 [SD 4·26] vs 8·67 [2·66]; Cohen's d=1·87).
Interpretation:
Our findings establish LLMs as promising tools for modelling affective processes relevant to human psychopathology. By reproducing key psychological phenomena, LLMs might enable the experimental investigation of mechanisms underlying mental disorders and facilitate the preliminary screening of novel therapeutic interventions, potentially accelerating progress in a field historically constrained by the scarcity of effective model systems.
Funding:
None.
Related Concept Videos
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in situations...
Typical Model Studies
Clearance Models: Physiological Models
The organ's clearance rate depends on the blood flow to the organ and the extraction ratio (E). The extraction ratio describes the organ's proficiency in drug...
Mechanistic Models: Compartment Models in Individual and Population Analysis
Theoretical Approaches to Psychological Disorder
Biological approach
The biological approach posits that internal, organic factors are the primary causes of such disorders. This perspective emphasizes brain structure and function, genetic predispositions, and neurotransmitter imbalances. For example, schizophrenia has been associated with both genetic...
Behavioral Genetics and Its Designs
The primary methodologies used in behavior genetics include family studies, twin studies, and adoption studies, each providing unique...