Related Experiment Videos
Data-centric review of multimodal emotion recognition: datasets, feature extraction, data fusion and limitations
Manisha Khanra1, Sanjiban Sekhar Roy1
1Vellore Institute of Technology, School of Computer Science and Engineering, Vellore, Tamilnadu, India.
Abstract:
Emotion recognition is a major challenge in affective computing. It enables systems to determine human emotions using multiple modalities, such as text, audio, facial expressions, and physiological data. The field of multimodal emotion recognition (MER) has achieved significant advances in recent years. Present studies tend to focus on single structures and fusion methods only. It provides little interaction between the modalities, input features, and learning models. An overview of MER studies that can be both modality-based and data-driven is provided in the present investigation. We present a methodical assesment of uni-, bi-, and multimodal frameworks, incorporating models to express emotions, feature extraction processes, fusion methods, learning methods, and benchmark datasets. We also provide an individual review approach that integrates performance with textual, acoustic, optical, and physical dimensions. Furthermore, it includes a thorough review for data-centric issues and mitigating approaches. We also consider benchmark multimodal datasets to tackle major problems such as modal diversity, annotating cost, data unavailability, and class imbalance. The MER workflow incorporates novel concepts, including linguistic models based on self-supervised learning. This review identifies opportunities for more robust and sustainable MER frameworks and highlights the remaining research challenges.