Related Experiment Videos
Speech Depression Screening via Multi-Scale Feature Enhancement and Emotion-Aware Contrastive Learning
Zhengyuan Chen1, Meihong Wu1,2
1School of Informatics, Xiamen University, Xiamen 361102, China.
Bioengineering (Basel, Switzerland)
|July 28, 2026
Summary
This study introduces a novel speech-based depression screening framework using multi-scale acoustic enhancement and contrastive learning. While improving specificity, the preliminary model needs further validation for accurate depression detection in short speech segments.
Area of Science:
- Computational Linguistics
- Psychiatry
- Machine Learning
Background:
- Speech-based depression screening offers a non-invasive method but faces challenges with short audio segments.
- Weak acoustic cues, limited semantic context, and overlapping participant data hinder accuracy.
Purpose of the Study:
- To develop a preliminary speech-based depression screening framework for short audio segments.
- To enhance acoustic feature extraction and representation calibration using advanced machine learning techniques.
Main Methods:
- Utilized a self-supervised Wav2Vec 2.0 model integrated with a Multi-Scale Convolution (MSC) module for acoustic feature enhancement.
- Implemented a Reference-Enhanced Contrastive Learning (ReCLR) mechanism with external emotion features for representation calibration.
- Incorporated Whisper-based transcription features to assess the utility of semantic information.
Main Results:
- Achieved an accuracy of 0.7021 and a specificity of 0.8485 for non-depressed participants in the DAIC-WOZ dataset.
- Demonstrated a sensitivity of 0.3571 for depressed participants, indicating room for improvement.
- The framework showed potential for specificity-oriented screening but requires further calibration.
Conclusions:
- Multi-scale acoustic enhancement and contrastive calibration show promise for improving specificity in short-segment screening.
- The current framework is preliminary, with limited sensitivity for depressed individuals.
- Further research is needed for threshold calibration, training stability, and external validation before clinical application.