A survey on deep multimodal learning for computer vision: advances, trends, applications, and datasets

Khaled Bayoudh1, Raja Knani2, Fayçal Hamdaoui3

  • 1Electrical Department, National Engineering School of Monastir (ENIM), Laboratory of Electronics and Micro-electronics (LR99ES30), Faculty of Sciences of Monastir (FSM), University of Monastir, Monastir, Tunisia.

The Visual Computer
|June 16, 2021
PubMed
Summary

Deep multimodal learning integrates diverse data types like images and text for computer vision. This review explores key concepts, fusion techniques, and future research directions in this rapidly advancing field.