Multimodal Sensing for Depression Risk Detection: Integrating Audio, Video, and Text Data

Zhenwei Zhang1,2, Shengming Zhang3, Dong Ni1,2

  • 1School of Biomedical Engineering, Health Science Center, Shenzhen University, Shenzhen 518060, China.

PubMed
Summary

This study introduces a novel Audio, Video, and Text Fusion-Three Branch Network (AVTF-TBN) for objective depression risk detection. The multimodal deep learning model effectively fuses sensor data, improving diagnostic accuracy.