Combining a parallel 2D CNN with a self-attention Dilated Residual Network for CTC-based discrete speech emotion

Ziping Zhao1, Qifei Li1, Zixing Zhang2

  • 1College of Computer and Information Engineering, Tianjin Normal University, Tianjin, China.

Summary

This study introduces an efficient deep neural network for speech emotion recognition (SER), utilizing parallel convolutional layers and self-attention networks with Connectionist Temporal Classification loss. The novel hybrid architecture improves SER performance by effectively modeling long temporal contexts in speech.

Related Concept Videos