Related Experiment Videos
On-premises open-source large language models for privacy-preserving multimodal depression screening
Soonjun Kwon1, Yihyun Kim1, Min Jhon2
1Department of Biomedical Informatics, Korea University College of Medicine, Seoul, the Republic of Korea.
Objective:
Narrative and speech data can provide valuable signals for depression screening, yet privacy and data-governance requirements often limit the use of closed-based models in clinical practice. In addition, existing large language model (LLM)-based approaches are largely text-centric, and multimodal integration of acoustic features and structured clinical variables remains limited. This study aimed to develop and externally validate a privacy-preserving multimodal depression screening prediction framework using open-source large language models that integrate sociodemographic information, emotion-memory narratives, and acoustic features.
Method:
This study analyzed 3536 participants collected at Chonnam National University Hospital. Inputs combined sociodemographic and lifestyle variables, Korean transcripts of happy- and sad-memory narratives, and speech-derived extended Geneva minimalistic acoustic parameter set (eGeMAPS) features. To maintain prompt conciseness, statistically significant features were selected from the 88 eGeMAPS features extracted for each happy- and sad-memory narrative, with Mann-Whitney U tests conducted exclusively on the internal training split to prevent data leakage. Five open-source LLMs (Gemma-3-27B, Qwen-3-32B, Llama-3.3-70B, Phi4-14B, and gpt-oss-20b) were evaluated under zero-shot prompting, Chain-of-Thought prompting, and supervised fine-tuning. External validation used Extended Distress Analysis Interview Corpus (E-DAIC) (N = 275).
Results:
Under zero-shot prompting, the best internal F1-score was 0.735 (Gemma-3-27B). Chain-of-Thought prompting improved Llama-3.3-70B (F1-score = 0.708) but reduced performance for other models. Supervised fine-tuning improved all models, yielding internal accuracies of 0.852 to 0.881 and F1-scores of 0.818 to 0.865 across five models, corresponding to F1 gains of 0.12 to 0.30 versus prompting-only approaches. In external validation, accuracy ranged from 0.764 to 0.822 and F1-score ranged from 0.683 to 0.807.
Conclusion:
This study suggests that multimodal open-source LLMs integrating clinical variables, narrative text, and acoustic features can support privacy-preserving depression screening in an on-premises setting. Supervised fine-tuning provided the most consistent performance improvements, and external validation supported robustness beyond the development cohort.
Related Concept Videos
Depressive Disorders: MDD and Dysthymia
Long-term Depression
Calcium Ion Concentration Mechanism
If over time, all...
Long-term Depression
Statistical Software for Data Analysis and Clinical Trials
Depression: Overview