Related Experiment Video
Updated: Sep 17, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
A Smartphone-Based Acoustic Machine Learning Pipeline for Detecting Suicidal Ideation: Case-Control Model Development
Min Lyu1,2, Lixin Tan2, Jun Xiao3
1Department of Medical Psychology, Army Medical University, No. 30 Gaotanyan Main Street, Shapingba District, Chongqing, 400038, China, +86-16623460852.
Background:
Suicidal ideation (SI) among university students is a growing public health concern. Self-report screening can be limited by concealment and delayed disclosure. We evaluated a leakage-resistant, proof-of-concept pipeline to detect SI from standardized smartphone-recorded speech.
Objective:
This study aimed to extract acoustic markers from brief smartphone-based reading tasks and develop machine learning models for suicide risk prediction in university students, enabling low-cost, scalable early screening to support campus mental health services.
Methods:
Questionnaire data and speech recordings were collected via a WeChat mini program. After screening and clinical confirmation, 96 participants (n=48, 50% with SI; n=48, 50% controls) were included. Age and sex were evaluated as potential confounders. Each participant read 16 standardized sentences. Acoustic features were extracted using openSMILE (version 3.0.2), yielding a 570D feature vector per utterance. To prevent leakage from multiple recordings per speaker, we used participant-level 5-fold cross-validation, assigning all recordings from each participant to a single fold. Seven machine learning algorithms were evaluated using area under the curve (AUC), accuracy, and F1-score.
Results:
Acoustic-based models discriminated participants with SI from control participants. The SI group was significantly older than the control group (P=.001). Random forest achieved an AUC of 0.813 (accuracy=0.748), and naive Bayes achieved an AUC of 0.806 (accuracy=0.757). Feature families related to pitch, mel-frequency cepstral coefficients, and harmonicity contributed to model performance.
Conclusions:
Standardized read speech captured via smartphones shows preliminary feasibility for SI discrimination under a leakage-aware evaluation design. External validation and testing with more naturalistic speech are warranted.
