Related Experiment Video
Updated: Aug 6, 2026

05:19
Closed-Loop Neurostimulation for Biomarker-Driven, Personalized Treatment of Major Depressive Disorder
Published on: July 7, 2023
Hierarchical Self-Supervised Representation Learning for Depression Detection from Speech
IEEE Journal of Biomedical and Health Informatics
|July 21, 2026
Summary
This study introduces HAREN-CTC, a novel framework for speech-based depression detection (SDD). By modeling acoustic-semantic interactions, it significantly improves the accuracy of identifying depression from speech patterns.
Area of Science:
- Computational linguistics
- Artificial intelligence in healthcare
- Psychiatric diagnostics
Background:
- Speech-based depression detection (SDD) is a promising non-invasive tool, but faces challenges due to subtle, heterogeneous, and temporally inconsistent depression markers.
- Current self-supervised learning (SSL) methods for SDD often oversimplify speech representations by using single layers or weighted sums, hindering the capture of acoustic-semantic interactions.
Purpose of the Study:
- To develop a hierarchical framework, HAREN-CTC, that effectively integrates multi-layer SSL speech representations for improved depression detection.
- To enhance the model's ability to capture interactions between low-level acoustic features and higher-level semantic context in speech.
- To address the temporal sparsity and irregularity of depression markers using an auxiliary Connectionist Temporal Classification (CTC) objective.
Main Methods:
- Proposed HAREN-CTC, a hierarchical framework utilizing two complementary SSL representation streams.
- Implemented semantic-conditioned cross-attention to emphasize context-informative acoustic patterns.
- Introduced an auxiliary CTC objective for weak alignment of temporal token representations with HuBERT-derived pseudo-tokens.
Main Results:
- HAREN-CTC outperformed strong baselines on the DAIC-WOZ and MODMA datasets, achieving Macro F1 scores of 0.81 and 0.82 in fixed-split settings.
- The model demonstrated consistent performance gains across most metrics under subject-level cross-validation.
- Results indicate the benefit of modeling acoustic-semantic interactions for robust speech-based depression assessment.
Conclusions:
- HAREN-CTC effectively models acoustic-semantic interactions within speech, leading to improved performance in speech-based depression detection.
- The proposed hierarchical framework and auxiliary CTC objective enhance robustness against the challenges of subtle and temporally irregular depression markers.
- This approach offers a more sophisticated method for leveraging SSL representations in clinical applications like depression assessment.