Related Experiment Videos
Speech Depression Screening via Multi-Scale Feature Enhancement and Emotion-Aware Contrastive Learning
Zhengyuan Chen1, Meihong Wu1,2
1School of Informatics, Xiamen University, Xiamen 361102, China.
Abstract:
Speech-based depression screening is a promising non-invasive approach, but short spontaneous-speech segments often contain weak acoustic cues, fragmented semantic context, and overlapping representations between depressed and non-depressed participants. This study addresses this problem by proposing a preliminary short-segment acoustic screening framework that integrates multi-scale feature enhancement and reference-enhanced contrastive calibration. Based on the self-supervised pretrained Wav2Vec 2.0 model, a Multi-Scale Convolution (MSC) module captures short-term vocal fluctuations and longer-range prosodic patterns. A Reference-Enhanced Contrastive Learning (ReCLR) mechanism further calibrates latent acoustic representations using external emotion-reference features. Whisper-based transcription features are also incorporated to test whether short-segment semantic information can complement acoustic cues. In the participant-independent DAIC-WOZ evaluation, the proposed framework achieved an accuracy of 0.7021, a specificity of 0.8485 for non-depressed participants, and a sensitivity of 0.3571 for depressed participants. These results indicate that multi-scale acoustic enhancement and emotion-reference contrastive calibration may improve specificity-oriented short-segment screening, but the limited depressed-class sensitivity shows that the framework remains preliminary and requires stronger threshold calibration, training stability, and external validation before practical deployment.