Related Experiment Videos
Spatial-Temporal Relation Enhancement for Speech Emotion Recognition from Acoustic Signals Using Fibonacci Encoding
Shicong Huang1, Jinghao Zhang1, Jingchao Xu1
1Sydney Smart Technology College, Northeastern University at Qinhuangdao, Qinhuangdao 066004, China.
Abstract:
Speech emotion recognition (SER) infers affective states from speech signals, but positional encoding for acoustic tokens remains underexplored in Transformer-based SER. Existing models often reuse encodings designed for text and do not explicitly account for the different sequential and two-dimensional structures of acoustic representations. We propose Fibonacci Position Embedding (FPE) and Fibonacci Target Shutter (FTS). FTS constructs overlapping candidate-index sets over time-frequency token grids, and FPE samples a Fibonacci index and applies its modulo-wrapped, dimension-dependent phase rotation to query and key vectors. The modules are integrated into STRE-Former, which fuses Wav2Vec, log-mel spectrogram, and MFCC representations through asymmetric cross-representation attention with representation-specific positional encodings. We also introduce an implementation-consistent conditional-entropy formulation that quantifies uncertainty in recovering a token location from its sampled positional representation; this quantity characterizes positional ambiguity rather than downstream modeling capacity. Experiments over 64 positional-encoding combinations on IEMOCAP and MELD identify dataset-dependent highest-observed configurations, reaching 74.21% weighted accuracy on IEMOCAP-4, 74.54% on IEMOCAP-6, and 49.44% on MELD. These empirical observations suggest that the relative behavior of positional-encoding strategies may depend on the acoustic representation and evaluation dataset, rather than supporting a single universally optimal scheme.