Related Experiment Video
Updated: Jun 28, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.5K
A Privacy-Preserving Unsupervised Speaker Disentanglement Method for Depression Detection from Speech
Vijay Ravi1, Jinhan Wang1, Jonathan Flint2
1Department of Electrical and Computer Engineering, University of California Los Angeles, California, USA 90095.
Summary
This study introduces an unsupervised method for speaker disentanglement in speech-based depression detection, enhancing patient privacy. The novel approach improves depression detection accuracy while effectively masking speaker identity without needing speaker labels.
Area of Science:
- Computational linguistics
- Machine learning
- Speech processing
Background:
- Current depression detection from speech often requires speaker labels, leading to privacy concerns.
- Existing methods can be unstable and add complexity through adversarial domain prediction.
Purpose of the Study:
- To develop an unsupervised speaker disentanglement method for privacy-preserving depression detection.
- To improve the accuracy and reduce the complexity of speech-based depression detection systems.
Main Methods:
- An unsupervised approach reducing cosine similarity between latent spaces of depression and speaker classification models.
- Utilizing ComparE16 features and an LSTM-only model on the DAIC-WOZ dataset.
- Score-level fusion with a Word2vec-based text approach.
Main Results:
- Achieved an F1-Score of 0.776 and a speaker de-identification (DeID) score of 92.87%, outperforming adversarial methods.
- Demonstrated superior performance compared to an adversarial counterpart (F1-Score 0.762, DeID 68.37%).
- Score-level fusion enhanced performance to an F1-Score of 0.830.
Conclusions:
- The proposed unsupervised method enhances privacy in speech-based depression detection without speaker labels.
- This approach reduces model complexity and improves performance over baseline and adversarial methods.
- Speaker disentanglement is complementary to text-based methods, offering significant performance gains when combined.

