Related Experiment Video
Updated: Nov 2, 2025

03:31
Author Spotlight: Enhancement of Salient Object Detection for Smart Grid Applications
Published on: December 15, 2023
730
A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets
Khaled Bayoudh1, Raja Knani2, Fayçal Hamdaoui3
1Electrical Department, National Engineering School of Monastir (ENIM), Laboratory of Electronics and Micro-electronics (LR99ES30), Faculty of Sciences of Monastir (FSM), University of Monastir, Monastir, Tunisia.
Summary
Deep multimodal learning integrates diverse data types like images and text for computer vision. This review explores key concepts, fusion techniques, and future research directions in this rapidly advancing field.
Area of Science:
- Computer Vision
- Artificial Intelligence
- Machine Learning
Background:
- Multimodal learning has seen significant growth, particularly in computer vision.
- Deep learning algorithms enable processing of diverse data streams (modalities) like visual and textual content.
- Extracting patterns from unstructured, real-world multimodal data remains a key research goal.
Purpose of the Study:
- To enhance understanding of deep multimodal learning for the computer vision community.
- To explore the generation of deep models integrating heterogeneous visual cues across sensory modalities.
- To provide a comprehensive overview of current research, applications, and datasets.
Main Methods:
- Surveying six key perspectives: multimodal data representation, fusion (traditional and deep learning-based), multitask learning, alignment, transfer learning, and zero-shot learning.
- Reviewing current multimodal applications in computer vision.
- Compiling benchmark datasets for various vision domains.
Main Results:
- A structured summary of six core concepts in deep multimodal learning.
- An overview of existing multimodal applications and relevant datasets.
- Identification of current limitations and challenges in the field.
Conclusions:
- Deep multimodal learning offers powerful tools for analyzing complex, real-world data.
- Further research is needed to address existing challenges and unlock future potential.
- This work provides a roadmap for researchers in deep multimodal learning for computer vision.
