Related Experiment Videos
Advancing Generative Artificial Intelligence for Clinical Assessment: Commentary on Chen, Horga, and Escola 2026
George Chatzisofroniou1, Alex S Cohen1,2
1Department of Psychology, Louisiana State University, Baton Rouge, LA 70803, United States.
Abstract:
Chen et al. provide proof-of-concept evidence that a general-purpose LLM, using zero-shot training, can derive clinical ratings from transcripts of semi-structured interviews in the AMP-SCZ. In this commentary, we complement this line of research with three considerations. First, the psychometric evaluation strategy would benefit from explicit "use case" specifications, since acceptable performance thresholds differ substantially across various uses. In the Chen et al., data, there is relatively compressed symptom-severity distribution in the data that may obscure its reliability and validity for some, but not all, use cases. Second, defining benchmarks requires consideration. The reliability benchmark used in this study are defensible but may not be ideal for CHR populations and may not fully account for information asymmetries (e.g., evaluating nonverbal behavior from transcripts) between human and LLM raters. Involving consensus or multi-rater benchmarks might better establish ground truth and help evaluate the scientific value of these technologies. Third, involves reproducibility, as single-prompt and single-system validation leaves open how results may be idiosyncratically tied to a specific prompt wording of a specific LLM model. This is a particular issue given the rapid deprecation and change associated with LLM models. To advance the authors ambitious and important research agenda, we recommend fixed annotated validation sets with consensus benchmarks tied to pre-specified use-cases and validation protocols. Additionally, multi-model and ensemble approaches may be important for generalizing results beyond specific LLM technologies. Addressing these issues would strengthen confidence in deploying LLM-derived ratings as a scalable, multilingual tool for schizophrenia research and clinical practice.