Related Experiment Video
Updated: Jun 7, 2025

Author Spotlight: Investigating the Impact of Emotional Prosodies on Voice Recognition and Perception
Published on: August 9, 2024
Speech based suicide risk recognition for crisis intervention hotlines using explainable multi-task learning
Zhong Ding1, Yang Zhou2, An-Jie Dai3
1Psychological Science and Health Research Center, China University of Geosciences, Lumo Road, Wuhan 430074, Hubei, China; Institute of Education, China University of Geosciences, Lumo Road, Wuhan 430074, Hubei, China; School of Automation, China University of Geosciences, Lumo Road, Wuhan 430074, Hubei, China.
Background:
Crisis Intervention Hotline can effectively reduce suicide risk, but suffer from low connectivity rates and untimely crisis response. By integrating speech signals and deep learning to assist in crisis assessment, it is expected to enhanced the effectiveness of crisis intervention hotlines.
Methods:
In this study, a crisis intervention hotline suicide risk speech dataset was constructed, and the speech was labeled based on the Modified Suicide Risk Scale. On the dataset, the variability of speech duration between different callers and different speech high-level features were explored across callers. Finally, this study proposed a data-theoretically dual-driven, gender-assisted speech crisis recognition method based on multi-tasking and deep learning, and the results of the model were obtained through five-fold cross-validation.
Results:
Analysis of the dataset demonstrated gender differences in callers, with male callers speaking more in crisis calls compared to females. Feature analysis revealed significant differences between crisis callers in terms of emotional intensity of speech, speech rate and texture. The proposed method outperformed other methods with an F1 score of 96 % on the validation data, and feature visualization of the model also demonstrated the validity of the method.
Limitations:
The sample size of this study was limited and ignored information from other modalities.
Conclusion:
These findings demonstrated the effectiveness of the proposed model in speech crisis recognition, and the statistical data analysis enhanced the Interpretability of the model, while showing that the integration of data and theoretical knowledge facilitates the effectiveness of the method.
More Related Videos
09:09Foreign Accent and Forensic Speaker Identification in Voice Lineups: The Influence of Acoustic Features Based on Prosody
Published on: September 27, 2024
07:31Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020