Related Experiment Video
Updated: Jul 27, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Attention guided learnable time-domain filterbanks for speech depression detection.
Wenju Yang1, Jiankang Liu1, Peng Cao1
1College of Computer Science and Engineering, Northeastern University, Shenyang, 110819, Liaoning, China; Key Laboratory of Intelligent Computing in Medical Image, Ministry of Education, Northeastern University, Shenyang, 110819, Liaoning, China.
This study introduces a novel speech depression detection (SDD) framework, DALF, which uses learnable filters to extract biologically meaningful acoustic features for improved depression screening. DALF significantly outperforms existing methods, offering a promising tool for early mental health detection.
Area of Science:
- Computational linguistics
- Psychiatry
- Machine learning
Background:
- Effective depression screening remains a challenge globally.
- Current speech depression detection (SDD) models often rely on fixed spectral features, limiting fine-grained analysis.
- There is a need for advanced methods to facilitate large-scale depression screening through speech analysis.
Purpose of the Study:
- To develop and evaluate a novel joint learning framework (DALF) for speech depression detection (SDD).
- To enable large-scale screening of depression by improving the accuracy and interpretability of SDD.
- To identify novel acoustic biomarkers for depression from speech signals.
Main Methods:
- Proposed a joint learning framework (DALF) incorporating depression filterbanks features learning (DFBL) and multi-scale spectral attention learning (MSSA).
- DFBL utilizes learnable time-domain filters to extract biologically meaningful acoustic features.
- MSSA guides filters to focus on relevant frequency sub-bands, enhancing feature representation.
- Introduced a new dataset, Neutral Reading-based Audio Corpus (NRAC), for depression analysis.
Main Results:
- The DALF model achieved state-of-the-art performance, with an F1 score of 78.4% on the DAIC-woz dataset.
- On the NRAC dataset, DALF achieved high F1 scores of 87.3% and 81.7% on its two parts.
- Analysis revealed a key frequency range (600-700Hz) as a potential acoustic biomarker for SDD.
Conclusions:
- The DALF model offers a promising and interpretable approach for speech depression detection.
- Learnable time-domain filters and multi-scale spectral attention can effectively capture depression-related speech characteristics.
- The identified frequency range provides valuable insights for developing more effective depression screening tools.
More Related Videos
05:19Author Spotlight: Therapeutic Benefit of Closed-Loop Deep Brain Stimulation in Depression Treatment
Published on: July 7, 2023
08:25Combined Invasive Subcortical and Non-invasive Surface Neurophysiological Recordings for the Assessment of Cognitive and Emotional Functions in Humans
Published on: May 19, 2016
Related Concept Videos
Classification of Signals
A continuous-time signal holds a value at every instant in time, representing information seamlessly. In contrast, a discrete-time signal holds values only at specific moments, often denoted as x(n), where...
Long-term Depression