Related Experiment Video
Updated: Jun 6, 2025

05:48
Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
1.4K
Multimodal machine learning for language and speech markers identification in mental health
Georgios Drougkas1, Erwin M Bakker1, Marco Spruit2,3
1Leiden Institute of Advanced Computer Science, Leiden University, Leiden, The Netherlands.
BMC Medical Informatics and Decision Making
|November 23, 2024
Summary
This study shows that multimodal approaches, combining text and audio data, can outperform unimodal methods for diagnosing a wide range of mental health disorders, particularly in identifying positive cases. Further refinements could enhance multimodal diagnostic accuracy.
Area of Science:
- Computational psychiatry
- Machine learning in healthcare
- Multimodal data analysis
Background:
- Existing research often uses unimodal approaches for diverse disorders or multimodal for single disorders.
- This study addresses the gap by compiling markers for various mental illnesses for both unimodal and multimodal diagnostics.
Purpose of the Study:
- To determine if multimodal approaches outperform unimodal methods in diagnosing a wide range of mental health disorders.
- To evaluate the effectiveness of combining text and audio data for mental health diagnostics.
Main Methods:
- Utilized the DAIC-WOZ dataset, focusing on text and audio modalities.
- Developed unimodal text and audio models, and a multimodal model using early fusion.
- Applied machine learning algorithms (SVM, Logistic Regression, Random Forests, Dense Layers) and evaluated using accuracy, AUC-ROC, and F1 scores.
Main Results:
- Unimodal text models achieved 78-87% accuracy and 85-93% AUC-ROC.
- Unimodal audio models achieved 64-72% accuracy and 53-75% AUC-ROC.
- Multimodal models showed comparable accuracy (80-87%) and AUC-ROC (84-93%) to text models, but superior F1 scores, especially for positive cases.
Conclusions:
- Multimodal models demonstrated potential to outperform unimodal approaches, particularly with refined feature engineering and label creation.
- Highlights the significance of multimodal integration for advancing mental health diagnostics.
- Suggests future research directions in advanced fusion techniques and deep learning models.

