Related Experiment Video
Updated: Jan 11, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Depression screening with textual and audio features based on large language models and machine learning
Yu Jin1, Xin Chen1, Xintian Hong2
1Department of Statistics, Faculty of Arts and Sciences, Beijing Normal University, Beijing, China.
Background:
Depression is a complex disorder that cannot be fully screened by textual features alone, as audio features capture additional psychomotor and affective changes. This study integrates textual and audio features for depression screening and compares the performance of various machine learning models.
Methods:
This study used a large-scale, multimodal psychology dataset of 1275 participants (707 males, 568 females; aged 12-16 years) that integrates PHQ-9 scores, textual interview responses, and mel-spectrograms derived from audio recordings. Textual features were calculated using suicide risk scores from the Chinese Suicide Dictionary (CSD), emotional polarity probabilities, and depression severity probabilities generated by large language models (LLMs). For audio data, we estimated the combination of emotion status by the frequency (ratio) of eight emotions, which applied a fine-tuned U-Net model with mel spectrograms, mel-frequency cepstral coefficients (MFCCs), and chroma features. Finally, these features were combined and evaluated with five machine learning models using eight metrics to identify the best-performing model.
Results:
Among the five machine learning methods, multimodal fusion outperformed unimodal approaches (text-only and audio-only) with the lowest MAE and RMSE. The RFR model showed the best performance for depression prediction (Accuracy = 0.98 and Precision = 0.98) with the combination of prompt3 from LLMs. The most important features for depression prediction were depression severity, negative and positive emotional polarity, and suicide risk from textual features, and emotional features (happy, angry, neutral, and surprise) from audio features.
Conclusions:
Combining audio and textual features improved depression screening accuracy. Future research could include facial expressions and physiological indicators to further enhance screening performance.
Related Concept Videos
Long-term Depression
Calcium Ion Concentration Mechanism
If over...
Long-term Depression
Depression: Overview
Depressive Disorders: MDD and Dysthymia
Depressive Disorders: Etiology
Biological Factors in Depression
Biological predispositions significantly influence the risk of developing depressive disorders. Genetic studies highlight the role of variations in the serotonin transporter...
