Related Experiment Video
Updated: Jul 3, 2026

Integrating Computerized Linguistic and Social Network Analyses to Capture Addiction Recovery Capital in an Online Community
Published on: May 31, 2019
Beyond text: Using network centralities and AI models to detect suicide risk on Reddit
Golnaz Nikmehr1, Aritz Bilbao-Jayo1, Aitor Almeida1
1Deustotech, University of Deusto, Bilbao, Spain.
None:
Suicide prevention research increasingly leverages social media data, where individuals often share personal experiences and emotions. While most studies rely on textual analysis to detect suicidal ideation, the structural characteristics of online social networks remain underexplored. In this exploratory pilot-scale study, we investigate whether network centrality metrics derived from Reddit interaction graphs can provide complementary signals for suicide risk detection beyond text-centric approaches. We construct directed user interaction graphs and extract 14 centrality metrics, including degree, betweenness, eigenvector, authority, and Katz centralities, to evaluate their predictive power. Using traditional machine learning classifiers, we assess the performance of these graph-based features in distinguishing suicidal from non-suicidal users. To further explore the interplay between network structure and textual content, we integrate centrality features with Sentence-BERT (SBERT) embeddings within a Graph Neural Network (GNN) framework. Within the constraints of a small-scale dataset, the best-performing GNN with SBERT achieved 67% accuracy and a 67% F1-score, representing a marginal improvement over the Random Forest baseline (64% accuracy; F1 = 66%). Results show that network centralities alone provide meaningful signals of suicidal risk, revealing patterns of isolation, influence, and social connectedness that characterize vulnerable users. Although these findings provide preliminary evidence for the complementary value of structural network analysis, they should be interpreted cautiously due to dataset scale and keyword-based sampling constraints.