Related Experiment Videos
Emotion recognition from body movement through interpretable motion-aware sequential modeling
Sergio Esteban-Romero1, Iván Martín-Fernández1, Rubén San-Segundo1
1Grupo de Tecnología del Habla y Aprendizaje Automático (THAU Group), Information Processing and Telecommunications Center, E.T.S.I. de Telecomunicación, Universidad Politécnica de Madrid (UPM), Madrid, Spain.
None:
Emotion recognition from bodily movement remains a challenging problem, particularly when only pose-based motion sequences are available and emotionally informative content is not uniformly distributed across time. In this work, we propose a Window Transformer architecture grounded in the Multiple Instance Learning (MIL) paradigm to address this challenge. Rather than processing the full sequence as a single temporal stream, the model decomposes it into overlapping windows and learns to assign greater relevance to those segments containing stronger emotional content. This formulation provides a more interpretable framework, since the learned relevance scores reveal which temporal regions drive the final prediction, while also yielding richer, context-aware representations of each segment. We evaluate the proposed approach on two publicly available datasets, MEED and DIEM-A, and compare it against a standalone Transformer baseline under different batch size and window configuration settings. The Window Transformer consistently outperforms the baseline and exhibits a more stable behavior across training configurations, achieving best accuracies of 56.35 ± 2.67% on MEED and 23.84 ± 0.83% on DIEM-A in a subject-independent scenario. Beyond performance gains, this work also establishes new benchmark values on both datasets, providing reference results that future research on bodily emotion recognition can build on.