AVaTER: Fusing Audio, Visual, and Textual Modalities Using Cross-Modal Attention for Emotion Recognition

Avishek Das1, Moumita Sen Sarma1, Mohammed Moshiul Hoque1

  • 1Department of Computer Science and Engineering, Chittagong University of Engineering and Technology, Chittagong 4349, Bangladesh.

Sensors (Basel, Switzerland)
|September 28, 2024
PubMed
Summary

Researchers developed a new multimodal Bangla dataset and framework for emotion recognition, improving accuracy by integrating audio, video, and text data.