Related Experiment Videos
Speech Depression Screening via Multi-Scale Feature Enhancement and Emotion-Aware Contrastive Learning
Zhengyuan Chen1, Meihong Wu1,2
1School of Informatics, Xiamen University, Xiamen 361102, China.
Bioengineering (Basel, Switzerland)
|July 28, 2026
Summary
This study introduces a novel speech-based depression screening framework using multi-scale acoustic enhancement and contrastive learning. While improving specificity, the preliminary model needs further validation for accurate depression detection.
Area of Science:
- Artificial Intelligence
- Computational Linguistics
- Psychiatry
Background:
- Speech-based depression screening offers a non-invasive method.
- Short speech segments present challenges due to weak acoustic cues and overlapping participant data.
Purpose of the Study:
- To develop a preliminary short-segment acoustic screening framework for depression.
- To enhance feature representation using multi-scale enhancement and contrastive calibration.
Main Methods:
- Utilized Wav2Vec 2.0 for self-supervised pretraining.
- Integrated a Multi-Scale Convolution (MSC) module for acoustic feature extraction.
- Employed Reference-Enhanced Contrastive Learning (ReCLR) with external emotion features.
- Incorporated Whisper-based transcription features for semantic context.
Main Results:
- Achieved 0.7021 accuracy and 0.8485 specificity for non-depressed participants on the DAIC-WOZ dataset.
- Demonstrated a sensitivity of 0.3571 for depressed participants.
- Indicated potential for specificity-oriented screening but highlighted limitations in depressed-class sensitivity.
Conclusions:
- Multi-scale acoustic enhancement and contrastive calibration show promise for improving specificity in short-segment screening.
- The framework is preliminary and requires further development, including threshold calibration, training stability, and external validation for clinical use.