Related Experiment Video
Updated: Aug 28, 2026

Memorization-Based Training and Testing Paradigm for Robust Vocal Identity Recognition in Expressive Speech Using Event-Related Potentials Analysis
Published on: August 9, 2024
Automatic analysis of speech representations to assess psychological distress
Sara Fernández-Velasco1, Jose Moreno-Mesa2, Daniel Escobar-Grisales2
1Telematics Department, Universidad del Cauca, Popayán, Colombia.
Background:
Current mental health diagnostic methods are limited by subjective clinical interpretation. Automatic speech analysis is a promising technology for objective assessment.
Objective:
To evaluate and compare different speech-based representations (acoustic, phonetic, and time-frequency) and deep learning-based embeddings for discriminating symptoms associated with psychological distress.
Methods:
A secondary analysis of the Distress Analysis Interview Corpus (DAIC-WOZ) was conducted using recordings from 125 participants (3,069 responses). Speech representations included phonation, articulation, and prosody features extracted with DisVoice; phonetic features extracted with Phonet; time-frequency representations derived from Mexican hat wavelets; and deep embeddings extracted with the multilingual Wav2Vec 2.0 model XLSR-53. Two classification strategies were addressed at the response and participant levels using a Fully Connected Neural Network (FCNN) and a Support Vector Machine (SVM), respectively.
Results:
Prosody at the participant level achieved the highest mean performance (F1-score 0.67 0.07; accuracy 0.64 0.10; AUC 0.65 0.12), followed by participant-level phonation (F1-score 0.59 0.16; accuracy 0.61 0.14; AUC 0.65 0.16). Conversely, participant-level aggregation of deep embeddings yielded lower performance (F1-score 0.48 0.19; accuracy 0.55 0.13; AUC 0.52 0.15), failing to surpass traditional features. Response-level performance remained close to chance. Phonet and wavelet representations did not improve performance over prosody or phonation.
Conclusion:
Participant-level analysis provided more robust and consistent discriminative patterns than response-level approaches. Prosody and phonation achieved the best performance across speech representations, while phonetic, time-frequency, and deep speech representations did not outperform the best acoustic baseline. These findings suggest that, within the evaluated experimental setting, the aggregation strategy appears to have a stronger influence on performance than increasing representational complexity.