Related Experiment Video
Updated: Jan 16, 2026

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
TSFNet: A Temporal-Spectral Fusion Network for advanced speech emotion recognition in medical applications
Xinran Li1, Peilin Huang1, Xiaojiang Peng1
1School of Artificial Intelligence, Shenzhen Technology University, 3002 Lantian Road, Shenzhen, 518118, Guangdong, China.
Abstract:
Speech emotion recognition (SER) is a critical component in enhancing communication systems and human-machine interaction, with significant potential for applications in the medical field. Although existing SER methods that combine temporal and spectral features have achieved notable advancements, they still encounter a big challenge in capturing emotional nuances, which are vital in medical diagnostics and patient care. In this study, we introduce a straightforward yet highly efficient network called TSFNet, which is the Temporal-Spectral Fusion Network via a Large-scale Pre-trained Model. This network is specifically designed to effectively process intricate emotional nuances by seamlessly integrating temporal and spectral information present in speech signals. By leveraging the capabilities of a large-scale pre-trained model, which serves as a powerful plug-and-play component for extracting and learning the temporal characteristics of speech, TSFNet enables a more accurate capture of complex emotional details crucial for medical applications. Extensive experiments are conducted on publicly available datasets, to evaluate the performance of TSFNet. Extensive experiments conducted on six public datasets demonstrate that TSFNet significantly outperforms existing baselines, achieving unweighted accuracies of 95.57% for Savee, 92.67% for Crema-D, 85.71% for IEMOCAP, 100.00% for Tess, 95.86% for Emovo, and 80.43% for Meld. It means that TSFNet has the potential in advancing medical diagnostic tools and patient monitoring systems.
