Feasibility of AI-powered assessment scoring: Can large language models replace human raters?

Michael Jaworski1, Jacob Balconi2, Celeste Santivasci3

  • 1Department of Psychology, Louisiana State University, Baton Rouge, LA, USA.

PubMed
Summary

ChatGPT-4.5 shows high accuracy in scoring the Brief International Cognitive Assessment for Multiple Sclerosis (BICAMS) before public release. However, its reliability decreased post-release, indicating potential for LLMs in neuropsychological assessment with optimization.