Related Experiment Video
Updated: Jun 5, 2025

08:45
Mapping Cortical Dynamics Using Simultaneous MEG/EEG and Anatomically-constrained Minimum-norm Estimates: an Auditory Attention Example
Published on: October 24, 2012
14.6K
Hierarchical transformer speech depression detection model research based on Dynamic window and Attention merge
Xiaoping Yue1, Chunna Zhang1, Zhijian Wang1
1School of Computer Science and Software Engineering, University of Science and Technology Liaoning, Anshan, Liaoning, China.
Peerj. Computer Science
|December 9, 2024
Summary
A new model, DWAM-Former, improves speech depression detection by effectively segmenting and integrating speech features. This approach enhances accuracy, outperforming previous methods in identifying depression from speech patterns.
Area of Science:
- Computational linguistics
- Affective computing
- Machine learning
Background:
- Speech depression detection is crucial for mental health assessment.
- Existing models struggle with segmenting and integrating speech data, leading to information loss.
- Hierarchical models offer potential but require refinement for optimal performance.
Purpose of the Study:
- To propose a novel Hierarchical Transformer model, DWAM-Former, for enhanced speech depression detection.
- To address challenges in segmenting and integrating depressed speech segments effectively.
- To reduce feature loss during information merging in speech analysis.
Main Methods:
- Developed DWAM-Former, a Hierarchical Transformer model incorporating dynamic window and attention merge.
- Introduced Learnable Speech Split module (LSSM) for phoneme and word segmentation.
- Implemented Adaptive Attention Merge (AAM) for representative feature generation and Variable-Length Residual module (VL-RM) to minimize feature loss.
Main Results:
- DWAM-Former achieved a competitive MF1 score of 0.788 on the DAIC-WOZ depression detection dataset.
- Demonstrated a 7.5% improvement over existing state-of-the-art methods.
- Effectively segmented and integrated speech features, preserving original information.
Conclusions:
- DWAM-Former offers a significant advancement in speech depression detection.
- The proposed modules (LSSM, AAM, VL-RM) effectively overcome limitations of previous models.
- This model shows promise for accurate and reliable mental health assessment through speech analysis.

