Related Experiment Video
Updated: May 29, 2025

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Suicide Risk Assessment on Social Media with Semi-Supervised Learning
Max Lovitt1,2, Haotian Ma1, Song Wang3
1Population Health Sciences, Weill Cornell Medicine, New York, New York, USA.
This study introduces a semi-supervised learning framework to improve automated suicide risk assessment from social media posts, addressing data limitations. The novel approach enhances model performance by utilizing pseudo-labeled data, aiding in early detection.
Area of Science:
- Computational linguistics
- Mental health informatics
- Machine learning for social good
Background:
- Social media platforms are increasingly used by individuals expressing suicidal ideation.
- Automated suicide risk assessment using natural language processing (NLP) is promising but hindered by insufficient labeled data and class imbalance.
- Existing methods struggle with the inherent challenges of real-world social media data for suicide risk detection.
Purpose of the Study:
- To develop a robust semi-supervised learning framework for automated suicide risk assessment.
- To address the challenges of limited labeled data and class imbalance in suicide-related social media content.
- To improve the accuracy and reliability of suicide risk detection models.
Main Methods:
- Proposed a semi-supervised framework combining 500 labeled and 1,500 unlabeled social media posts.
- Expanded the self-training algorithm with a novel pseudo-label acquisition process tailored for imbalanced datasets.
- Implemented manual verification of a subset of pseudo-labeled data to ensure quality and reliability.
- Evaluated multiple models, identifying RoBERTa as the optimal backbone for the framework.
Main Results:
- The semi-supervised framework significantly improved the model's ability to assess suicide risk.
- The novel pseudo-labeling strategy effectively handled class imbalance issues.
- Partially validated pseudo-labeled data, combined with ground-truth labels, enhanced predictive performance.
- RoBERTa demonstrated superior performance as the underlying model for risk assessment.
Conclusions:
- Semi-supervised learning, particularly with a refined pseudo-labeling process, is effective for suicide risk assessment on social media.
- The proposed framework offers a viable solution to data scarcity and imbalance problems in this critical domain.
- This approach holds potential for developing more accurate and scalable automated systems for suicide prevention.
More Related Videos
Related Concept Videos
Self-Help Support Groups
Accessibility and Cost-Effectiveness
One of the primary strengths of self-help...
Bullying
Self-Presentation: Self-Monitoring and Self-Handicapping
Modeling in Therapy
Participant Modeling
Participant modeling involves therapists demonstrating calm and effective behaviors in...
Social Anxiety Disorder

