Related Experiment Video
Updated: Aug 8, 2026

Implementation of a Real-Time Psychosis Risk Detection and Alerting System Based on Electronic Health Records using CogStack
Published on: May 15, 2020
Large language models can supplement the assessment of clinical high risk for psychosis
Luz Maria Alliende1, Rob Voigt2, Maeve Hoffman1
1Department of Psychology, Northwestern University.
Abstract:
Capturing psychosis risk before illness onset is an ongoing challenge for psychosis-spectrum studies. Natural language processing (NLP) tools can harness information embedded in the notes generated during clinical interviews to obtain more objective markers of psychosis risk; for this project, we used readily available assessor notes. The following project acts as proof of concept on the usefulness of AI-based tools to capture latent psychosis risk information. We used assessor notes for 2,077 Structured Interview for Psychosis-risk Syndromes interviews as input for different AI-based NLP models to produce metrics for psychosis risk, namely, a trained encoder-only model (ModernBERT), and zero-shot and few-shot instantiation of two decoder-only models (Large Language Model Meta AI Version 4 [LLaMA-4] and GPT-4.1 mini). AI-based risk metrics were compared with gold-standard ratings of clinical high risk for psychosis (CHR). The AI-based risk metrics' relationship to traditional psychosis risk scores (Shanghai-At-Risk-for-Psychosis [SHARP] and North American Prodrome Longitudinal Study [NAPLS]) was also assessed. Finally, we explored the added benefit of adding AI-based risk metrics to models predicting future participant conversion. All models performed above chance in classifying interview notes for the presence or absence of CHR syndromes. The trained encoder model performed the best out of all models determining the presence of CHR syndromes (accuracy = 82.67%, κ = .63). Positive CHR classification by the encoder model resulted in a 0.69 and 0.85 standard deviation increase in SHARP and NAPLS risk scores (p < .001). A 1 standard deviation increase in decoder-generated risk scores increased traditional risk scores between 0.24 and 0.47 standard deviations (all ps < .001). In exploratory analyses, decoder-generated risk scores incrementally improved models predicting conversion including SHARP but not NAPLS risk scores. Although our results need to be considered in the context of one consortium, AI-based NLPs show potential as an aid for the diagnosis of CHR syndromes, evaluating psychosis risk, and even predicting future conversion, even with suboptimal but readily available inputs (i.e., assessor notes). Future projects could use AI-based tools' potential in augmenting psychosis risk screenings and risk predictors. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
