Related Experiment Video
Updated: Aug 26, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Using Large Language Models to Measure Symptom Severity Scores in Patients At-Risk for Schizophrenia
Andrew X Chen1, Guillermo Horga1, Sean Escola1,2
1Department of Psychiatry, Columbia University/New York State Psychiatric Institute, New York, NY 10032, United States.
Background And Hypothesis:
Patients who are at clinical high risk (CHR) for schizophrenia need close monitoring of their symptoms to inform appropriate treatments. The Brief Psychiatric Rating Scale (BPRS) is a validated, commonly used research tool for measuring symptoms in patients with schizophrenia and other psychotic disorders; however, it is not commonly used in clinical practice as it requires a lengthy structured interview. We hypothesized we could utilize large language models (LLMs) to predict BPRS scores from clinical interview transcripts.
Study Design:
We used LLMs to predict BPRS scores from transcripts in 409 CHR patients from the Accelerating Medicines Partnership Schizophrenia cohort.
Study Results:
Despite the interviews not being specifically structured to measure the BPRS, the zero-shot performance of the LLM predictions compared to the true assessment (median concordance: 0.84, Intraclass Correlation Coefficient [ICC]: 0.73) approaches human inter- and intra-rater reliability. We further demonstrate that LLMs have expansive potential to improve and standardize the assessment of CHR patients via their accuracy in assessing the BPRS in foreign languages (median concordance: 0.88, ICC: 0.70), and integrating longitudinal information as one-shot or few-shot learning.
Conclusions:
LLMs may present a promising pathway to extract symptom severity scores from clinical interviews for improved monitoring of CHR patients.