Related Experiment Video
Updated: Sep 9, 2025

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
681
Feasibility of AI-powered assessment scoring: Can large language models replace human raters?
Michael Jaworski1, Jacob Balconi2, Celeste Santivasci3
1Department of Psychology, Louisiana State University, Baton Rouge, LA, USA.
The Clinical Neuropsychologist
|September 1, 2025
Summary
ChatGPT-4.5 shows high accuracy in scoring the Brief International Cognitive Assessment for Multiple Sclerosis (BICAMS) before public release. However, its reliability decreased post-release, indicating potential for LLMs in neuropsychological assessment with optimization.
Area of Science:
- Neuroscience
- Artificial Intelligence
- Medical Informatics
Background:
- Neuropsychological assessments like the Brief International Cognitive Assessment for Multiple Sclerosis (BICAMS) are crucial for monitoring disease progression.
- Manual scoring of these assessments is time-consuming and prone to human error.
- Large Language Models (LLMs) offer potential for automating such tasks.
Purpose of the Study:
- To evaluate the feasibility, accuracy, and reliability of using ChatGPT-4.5 for automated scoring of BICAMS protocols.
- To compare ChatGPT-4.5's scoring performance against human raters.
Main Methods:
- Thirty-five deidentified BICAMS protocols (including SDMT, CVLT-II, BVMT-R) were scored by two human raters and ChatGPT-4.5.
- Scoring involved uploading protocol scans and using structured prompts for ChatGPT-4.5.
- Interrater reliability, accuracy, and speed were assessed using ICCs, t-tests, and descriptive statistics.
Main Results:
- Before public release, ChatGPT-4.5 demonstrated strong interrater reliability with human raters (e.g., CVLT-II ICC = 0.992, SDMT ICC = 1.000).
- ChatGPT-4.5 completed scoring in under 9 minutes per protocol and identified errors missed by human raters.
- After public release, reliability decreased significantly (e.g., BVMT-R Trial 3 ICC = -0.046), and scoring discrepancies increased.
Conclusions:
- ChatGPT-4.5 showed comparable accuracy to human raters in scoring BICAMS protocols, particularly before its public release.
- Performance variability emerged after public release, highlighting the need for model and prompt optimization.
- LLMs hold promise for streamlining neuropsychological assessments, improving efficiency, and reducing errors, provided computational resources and optimization are adequate.

