Related Experiment Videos
Temporal and cross-site validation of an AI system for self-harm detection
Vlada Rozova1,2,3, Liuliu Chen2, Katrina Witt4,5
1Department of Data Science and AI, Monash University, Melbourne, Australia.
Abstract:
Adequate self-harm surveillance is a key part of suicide prevention. Our previous research demonstrated that an artificial intelligence (AI)-based system could effectively detect self-harm in emergency department triage notes. However, the system was developed using data from a single hospital, raising concerns about its generalisability. Here, we aim to validate the system prospectively and externally to better understand its portability across hospitals. We leveraged emergency department data from two Australian hospitals, with free-text triage notes manually annotated for self-harm. The AI system was developed using data from a major metropolitan hospital in Melbourne, spanning 2012-2017. The system combined extensive text normalisation with a self-harm classification model utilising 1931 selected features. We prospectively validated the model on 329,655 triage notes from the same hospital over the following four years. For external validation, we used 316,877 triage notes from 2012 to 2021 from a regional hospital 150 km outside Melbourne. On the test set, the model achieved an area under the precision-recall curve (PR AUC) of 0.84 with a 95% confidence interval of [0.82, 0.86]. This performance remained stable at the development site with a PR AUC of 0.84, 95% CI [0.83, 0.85]. When applied in the regional context, the model's ability to distinguish self-harm cases declined, resulting in an overall PR AUC of 0.78, 95% CI [0.77, 0.79]. While the text normalisation component was equally effective across the datasets, our analysis revealed that in regional settings, self-harm presentations are more likely to involve medication ingestion. At the metropolitan hospital, the AI system for self-harm detection maintained its epidemiological utility. At the regional hospital, the text normalisation process was effective, but performance was unstable primarily due to linguistic domain shift and differences in self-harm presentations.