Related Experiment Video
Updated: Jan 17, 2026

Author Spotlight: Deciphering Electrical Networks Behind Complex Brain Activities and Disorders
Published on: November 1, 2024
Neonatal pain speech emotion recognition based on Horizontal and Vertical Sparse Mask Time-Frequency Transformer
Jingjie Yan1, Wenjing Sun1, Boyan Sun1
1Nanjing University of Posts and Telecommunications, Nanjing 210003, China.
None:
Neonatal pain speech emotion recognition is a novel and challenging research topic that can effectively assist clinical medical staff in accurately determining the pain status of neonates. However, to date, few studies have addressed this topic, with the primary challenge being the lack of a neonatal pain speech database. Therefore, this paper firstly establishes a novel multi-category Neonatal Pain Speech (NPS) database. The NPS database in total contains 461 neonatal speech samples and covers the categories of severe pain, mild pain, crying, and calmness. Additionally, this paper also proposes a Horizontal and Vertical Sparse Mask Time-Frequency Transformer Network (HVSMTNet) for neonatal pain speech emotion recognition. HVSMTNet extracts speech features and inputs them into both a time-domain Transformer and a frequency-domain Transformer. Then, HVSMTNet presents a vertical and horizontal sparse mask mechanism into the QKV calculations of both the time-domain and frequency-domain Transformers to help the network assess the importance of speech segments in neonatal samples. The experimental results demonstrate the effectiveness of the NPS database, showing that it is feasible to distinguish between crying caused by pain and crying caused by non-pain reasons through speech. Moreover, compared to several mainstream speech emotion recognition methods, the proposed HVSMTNet model achieves the highest recognition rates in both experiments. In addition, we also conduct experiments with the HVSMTNet model on the CASIA database, further validating its good performance across different datasets.

