Related Experiment Videos
Structured Prototype Learning with Feature Fusion for Sparse and Asynchronous Audio-Visual Depression Recognition
Zhonghui Jin1, Pei He2, Yangming Guo3
1School of Software Engineering, Northwestern Polytechnical University, Xi'an 710072, China.
Abstract:
Audio-visual depression recognition in real-world scenarios is often challenged by temporal sparsity and cross-modal asynchrony, where depression-related cues may appear only in short segments and may not align precisely across modalities. Under such conditions, global pooling or dense attention tends to dilute sparse discriminative evidence with redundant context, leading to unstable utterance-level representations. To address this issue, we propose an audio-visual depression recognition framework that integrates modality feature adaptation, bidirectional cross-modal interaction, and graph-based prototype abstraction. Specifically, heterogeneous audio and visual streams are first transformed into compatible representations, after which bidirectional cross-attention models content-dependent dependencies across modalities without requiring index-wise correspondence. The fused tokens are then interpreted as graph nodes and aggregated into a compact set of semantic prototypes through graph convolution and differentiable soft clustering. In addition, audio perturbation is introduced during training as a task-oriented regularisation strategy for partial acoustic evidence loss and temporal misalignment. Experiments on the LMVD dataset demonstrate clear improvements on the primary depression-recognition task, while auxiliary evaluations on MIntRec and CMU-MOSI suggest that the structured prototype representation is beneficial for other temporally sparse audio-visual recognition tasks. These results indicate that structured prototype learning is effective for preserving sparse depression-related cues, while training-time perturbation provides a complementary regularisation effect under asynchronous multimodal conditions.