Related Experiment Video
Updated: Mar 31, 2026

Perceptual and Category Processing of the Uncanny Valley Hypothesis' Dimension of Human Likeness: Some Methodological Issues
Published on: June 3, 2013
Granularity paradox: how emotion taxonomies shape GPT-5's affective cognition and human-AI alignment
1School of Management, Guangdong University of Science and Technology, Dongguan, China.
Background:
Large Language Models (LLMs) have demonstrated exceptional capability in textual emotion detection. However, LLM evaluations often treat the emotion taxonomy-the "cognitive ruler" defining the emotional space-as a neutral background variable. The extent to which taxonomic complexity moderates LLM performance remains underexplored.
Method:
This study systematically evaluates the impact of emotion taxonomy on GPT-5's annotation behavior. We constructed a dataset of 2,848 Chinese Weibo posts. Five human annotators and GPT-5 (zero-shot) labeled the data across five distinct taxonomies, each with varying levels of granularity: SemEval (4 classes), Ekman (6 classes), Chinese SevenEmotions (7 classes), Plutchik (8 classes), and GoEmotions (27 classes). A rigorous experimental design, including randomized ordering and washout periods, was implemented to minimize sequence effects. By comparing the results of GPT-5 and manual annotation, the analysis is conducted across three dimensions: performance, consistency, and bias patterns.
Results:
Results reveal a significant "granularity paradox": GPT-5's performance is strongly negatively correlated with taxonomic complexity, with performance collapsing in fine-grained settings (GoEmotions). Crucially, we identified systematic misalignment mechanisms: (1) Consistency decay: Human-AI agreement significantly deteriorates as semantic boundaries blur in complex taxonomies; (2) Hyper-sensitivity bias: GPT-5 exhibits a tendency to over-interpret neutral texts as emotional, with false-positive rates increasing with taxonomy size; and (3) Arousal shift: The model consistently misclassifies low-arousal negative emotions (e.g., sadness) as high-arousal prototypes (e.g., fear/anger), reflecting a valence-based rather than nuance-based inference logic. Notably, the indigenous SevenEmotions did not yield superior cultural alignment compared to Western taxonomies.
Conclusion:
Our findings suggest that emotion taxonomies function as a critical hyperparameter that shapes the cognitive boundaries of GPT. While GPT shows promise, its reliability is compromised by complex taxonomies. Researchers must balance granular detail against model robustness when deploying LLMs for psychological analysis.
Related Concept Videos
The Influence of Affect on Cognition
The Influence of Cognition on Affect
Empathy
Stereotype Content Model
Cognitive Theories: Schachter-Singer Theory of Emotion
Physiological Arousal and Cognitive Labeling
According to this theory, when an individual experiences...
Cognitive Theories: Lazarus Mediational Theory of Emotion
Cognitive Appraisal and Emotional Response
Lazarus proposed that...

