Related Experiment Video
Updated: Jul 4, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Enhancing accuracy and privacy in speech-based depression detection through speaker disentanglement
Vijay Ravi1, Jinhan Wang1, Jonathan Flint2
1Department of Electrical and Computer Engineering, University of California, Los Angeles, CA, 90095, USA.
This study introduces novel methods to detect Major Depressive Disorder (MDD) using speech, while protecting patient privacy by disentangling speaker identity from depression signals. The approach improves depression detection accuracy and enhances voice privacy.
Area of Science:
- Artificial Intelligence
- Computational Linguistics
- Psychiatry
Background:
- Speech signals are increasingly recognized as valuable biomarkers for mental health assessment, particularly for the automatic detection of Major Depressive Disorder (MDD).
- Current methods often rely on speaker identity features, raising concerns about patient privacy and potential biases that can impair diagnostic accuracy.
- There is a need for advanced techniques that can extract depression-related information from speech without compromising sensitive speaker-specific data.
Purpose of the Study:
- To develop and evaluate novel methods for disentangling speaker-identity information from depression-related information in speech signals.
- To improve the accuracy of automatic Major Depressive Disorder (MDD) detection while simultaneously enhancing voice privacy.
- To overcome the limitations of existing approaches that over-rely on speaker identity features.
Main Methods:
- Proposed four distinct disentanglement methods: adversarial speaker identification (SID)-loss maximization (ADV), SID-loss equalization with variance (LEV), SID-loss equalization using Cross-Entropy (LECE), and SID-loss equalization using KL divergence (LEKLD).
- Conducted experiments using diverse input features (e.g., ComparE16, raw-audio) and model architectures on the DAIC-WOZ (English) and EATD (Mandarin) datasets.
- Quantified improvements in MDD detection performance (F1-Scores) and voice-privacy attributes using Gain in Voice Distinctiveness () and De-Identification Scores (DeID).
Main Results:
- The LECE method with ComparE16 features achieved state-of-the-art (SOTA) audio-only F1-Scores of 80% for MDD detection on the DAIC-WOZ dataset, with a of -1.1 dB and DeID of 85%.
- The ADV method with raw-audio signals attained an F1-Score of 72.38% on the EATD dataset, surpassing multi-modal SOTA, alongside a of -0.89 dB and DeID of 51.21%.
- All proposed disentanglement methods demonstrated improvements in both depression detection accuracy and voice-privacy metrics compared to baseline approaches.
Conclusions:
- Disentangling speaker-identity information from depression-related cues in speech is a viable strategy for privacy-preserving mental health assessment.
- The proposed methods offer a promising direction for developing robust and ethical speech-based depression detection systems.
- Reducing reliance on speaker identity features enhances system performance and safeguards sensitive patient information.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
05:22Dissociation of the Confounding Influences of Expectancy and Integrative Difficulty Residing in Anomalous Sentences in Event-related Potential Studies
Published on: May 9, 2019
Related Concept Videos
¹³C NMR: Distortionless Enhancement by Polarization Transfer (DEPT)
Deindividuation