Related Experiment Video
Updated: Jun 4, 2025

Author Spotlight: Addressing Technical and Subjective Challenges in Measuring Classroom Attention
Published on: December 15, 2023
Harnessing multimodal approaches for depression detection using large language models and facial expressions
Misha Sadeghi1, Robert Richer2, Bernhard Egger3
1Machine Learning and Data Analytics Lab (MaD Lab), Department Artificial Intelligence in Biomedical Engineering (AIBE), Friedrich-Alexander-Universität Erlangen-Nürnberg (FAU), Erlangen, 91052, Germany. misha.sadeghi@fau.de.
This study developed an automated method to predict depression severity using interview transcripts and facial data. Text-based analysis, enhanced with speech quality, proved most effective for accurate depression detection.
Area of Science:
- Computational psychiatry and mental health informatics.
- The intersection of natural language processing and automated depression detection.
- Machine learning applications in clinical diagnostic assessment.
Background:
Precise identification of depressive disorders remains a fundamental pillar of modern psychiatric diagnostics because accurate assessment is essential for determining the most effective treatment pathways for patients. Prior research has shown that clinical interviews provide rich datasets for evaluating psychological states, yet the manual interpretation of these interactions often introduces subjective bias into the diagnostic process. Traditional assessment methods frequently rely on the expertise of clinicians to synthesize verbal and non-verbal cues, which can lead to variability in severity ratings across different practitioners. Computational tools have recently begun to bridge the gap between qualitative observation and quantitative measurement by applying advanced algorithms to human behavioral data. Integrating diverse data streams such as speech patterns and visual expressions offers a promising path toward developing objective, scalable screening tools for mental health conditions. This absence of evidence motivated the development of a fully automated system capable of synthesizing linguistic and physiological indicators to predict depression severity with high precision.
Purpose Of The Study:
This research investigates a fully automated framework for predicting the severity of depressive symptoms by leveraging advanced computational techniques and diverse data modalities. The investigators sought to utilize the Extended-Distress Analysis Interview Corpus (E-DAIC) to refine the accuracy of diagnostic predictions through the application of sophisticated machine learning architectures. Utilizing Large Language Models (LLMs) allows for the extraction of nuanced depression-related indicators from interview transcripts that might be overlooked during standard clinical evaluations. The study aims to integrate facial data extracted from video frames with textual information to create a comprehensive multimodal model for severity prediction. Researchers intended to compare the relative efficacy of unimodal approaches, focusing on text-based features and facial features, against a combined multimodal experimental configuration. Establishing a robust baseline for severity prediction using the Patient Health Questionnaire-8 (PHQ-8) served as the primary benchmark for evaluating the performance of these automated models.
Main Methods:
The team utilized the Extended-Distress Analysis Interview Corpus (E-DAIC) as the primary data source for training and validating their automated depression detection system. Large Language Models (LLMs) processed the interview transcripts to identify and extract relevant psychological markers associated with the severity of depressive symptoms. Facial data extraction involved the systematic analysis of individual video frames to capture expressive variations and physiological cues that correlate with mental health status. The prediction model underwent rigorous training using the scores from the Patient Health Questionnaire-8 (PHQ-8) to ensure the outputs aligned with established clinical standards. Three distinct experimental configurations evaluated the performance of text-based features, facial features, and a combined multimodal approach to determine the most effective diagnostic strategy. Speech quality assessment served as an additional refinement layer for the textual data stream, providing a more nuanced understanding of the patient's communicative state during the interview.
Main Results:
Enhancing textual data with speech quality assessment yielded the most accurate diagnostic performance across all tested configurations in the automated depression detection framework. The optimized model achieved a mean absolute error (MAE) of 2.85, indicating a high level of precision in predicting the severity of depressive symptoms. Statistical evaluation revealed a root mean square error (RMSE) of 4.02, which further validates the reliability of the combined text and speech quality approach. Text-only models demonstrated significant robustness in predicting symptom severity, suggesting that linguistic content remains a powerful indicator of a patient's underlying mental state. Integrating facial features provided a secondary layer of data for the multimodal analysis, though these features were less predictive than the linguistic indicators on their own. Comparative testing indicated that the synthesis of linguistic features and speech quality metrics outperformed the facial-only and basic multimodal models in this specific study.
Conclusions:
Automated systems represent a viable path forward for enhancing mental health diagnostic workflows by providing objective and scalable tools for clinicians. The findings highlight the reliability of linguistic analysis and Large Language Models (LLMs) in identifying depressive indicators within the context of clinical interviews. Future research may focus on refining the integration of visual and auditory data streams to further improve the sensitivity and specificity of multimodal diagnostic tools. Implementing these computational models could streamline the initial screening process in clinical settings, allowing for more rapid identification of individuals requiring urgent psychiatric intervention. The study provides a foundation for developing more sophisticated diagnostic frameworks that can be applied to a wide range of mental health conditions beyond depression. These results suggest that large-scale automated screening is increasingly feasible, potentially expanding access to mental health assessments in underserved or remote populations.
Frequently Asked Questions
Based on this study's findings, Large Language Models (LLMs) extract depression-related indicators from interview transcripts by identifying linguistic patterns that correlate with the Patient Health Questionnaire-8 (PHQ-8) scores, enabling precise severity prediction.
The researchers found that the top-performing model, which combined textual data with speech quality assessment, achieved a mean absolute error (MAE) of 2.85 and a root mean square error (RMSE) of 4.02.
The E-DAIC dataset was utilized because it provides a comprehensive collection of interview transcripts and video frames, allowing the researchers to integrate textual features with facial data for a robust multimodal prediction framework.
The study's authors indicate that while text-only models are robust, the current system is specifically trained on the E-DAIC dataset and relies on the Patient Health Questionnaire-8 (PHQ-8) as the ground truth for severity.
The study's authors propose that automated depression detection using multimodal analysis paves the way for more efficient diagnostic tools that can be integrated into clinical workflows to support mental health professionals.
Related Concept Videos
Facial Feedback Hypothesis
Antidepressant Drugs: MAOIs and Other Agents
Depression: Overview

