Related Experiment Video
Updated: Sep 27, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.7K
Privacy-Preserving Deep Speaker Separation for Smartphone-Based Passive Speech Assessment.
Apiwat Ditthapron1, Emmanuel O Agu1, Adam C Lammert2
1Computer Science DepartmentWorcester Polytechnic Institute Worcester MA 01609 USA.
IEEE Open Journal of Engineering in Medicine and Biology
|April 11, 2022
Summary
Deep-MFCC bAsed SpeaKer Separation (Deep-MASKS) improves speech separation for health monitoring by reconstructing Mel-Frequency Cepstral Coefficients (MFCCs). This method enhances privacy and accuracy in analyzing speech impairments from smartphone recordings.
Area of Science:
- Speech processing
- Machine learning
- Biomedical engineering
Background:
- Smartphones offer passive speech impairment monitoring for conditions like Parkinson's disease and Alzheimer's.
- Speaker separation is vital for isolating target speech from background noise and cross-talk in recordings.
- Privacy concerns necessitate using speech features like Mel-Frequency Cepstral Coefficients (MFCCs) instead of raw audio.
Purpose of the Study:
- To introduce Deep MFCC bAsed SpeaKer Separation (Deep-MASKS) for improved speaker separation.
- To enable accurate speech analysis from smartphone recordings while preserving speaker privacy.
Main Methods:
- Deep-MASKS employs an autoencoder and Deep Neural Network (DNN) to reconstruct MFCCs.
- Speech representations (i-vector, x-vector, d-vector) are used for MFCC component reconstruction.
- The method processes continuous audio recordings, unlike utterance-based approaches.
Main Results:
- Deep-MASKS achieved up to a 44% reduction in Mean Squared Error (MSE) for MFCC reconstruction.
- The method reduced the additional bits needed for clean speech entropy by 36%.
- Performance surpassed existing baseline methods.
Conclusions:
- Deep-MASKS provides a privacy-preserving and accurate method for speaker separation using MFCCs.
- This technique is suitable for passive speech monitoring applications on smartphones.
- The DNN-based reconstruction offers superior performance over previous masking techniques.
Keywords:
Impact Statement—The proposed Deep-MASKS mitigates cross-talk in speech encoded as MFCC features, which are widely utilized to preserve voice privacy in passive health assessment and other speech applications on smartphonesMel-Frequency Cepstrum Coefficients (MFCCs)overlapped speechspeaker representationspeech separation
