Related Experiment Videos
Multi-Modal Depression Recognition Via Label-Aware Contrastive Learning and Cross-Modal Alignment
Abstract:
The diagnostic accuracy of depression is hindered by its heterogeneous presentation and the insufficiency of unimodal data. To address this, we introduce a novel framework, Multi-Modal Depression Recognition via Label-Aware Contrastive Learning and Cross-Modal Alignment (MMDR-LCA). Our core technical contribution is a cross-modal alignment mechanism that systematically models the dyadic relationships between text-audio, text-video, and audio-video modalities. Unlike conventional fusion, this mechanism explicitly mitigates latent modality conflicts, such as positive facial expressions deceptively masking negative verbal content, by extracting aligned and complementary emotional cues. Furthermore, this mechanism is integrated into a label-aware supervised contrastive learning framework. By leveraging continuous severity scores to construct soft-label contrastive losses, our approach bridges the semantic gap between symptom regression and binary diagnosis. This strategy encourages samples with similar psychological states to form closer representations, thereby mitigating the adverse effect of class imbalance in clinical datasets. Experimental results on the public E-DAIC benchmark, using the official training, validation, and test splits, show that MMDR-LCA achieves an MAE of 3.85 and an F1-score of 0.76. In addition, imbalance-aware analyses report sensitivity, specificity, and balanced accuracy, and a controlled class-balanced sampling experiment is included to examine the effect of the depressed/healthy imbalance. Comprehensive ablation studies further verify that cross-modal alignment helps handle inter-modal heterogeneity, while contrastive learning enhances discriminative decision boundaries and overall robustness. This work establishes a promising paradigm for AI-driven depression assessment.