Related Experiment Video
Updated: May 3, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
Published on: June 25, 2019
The answer may vary: large language model response patterns challenge their use in test item analysis.
1Department of Anesthesiology, Dartmouth Hitchcock Medical Center, Lebanon, NH, USA.
Large language models (LLMs) show limited ability to predict multiple-choice question (MCQ) performance metrics like difficulty and point biserial indices. Consistency of LLM responses is key for assessment development, not prediction of item characteristics.
Area of Science:
- Medical Education
- Artificial Intelligence in Assessment
Background:
- Validating multiple-choice questions (MCQs) requires extensive testing.
- Large language models (LLMs) offer potential for streamlining assessment development.
- Predicting psychometric properties of MCQs is a key challenge.
Purpose of the Study:
- To investigate LLMs' ability to predict MCQ difficulty and point biserial indices.
- To assess if LLMs can reduce the need for preliminary test population analysis.
- To compare LLM performance with human expert assessment.
Main Methods:
- Sixty anesthesiology MCQs were administered to five LLMs and clinical fellows.
- LLM response patterns, difficulty indices, and point biserial indices were analyzed.
- Spearman correlation coefficients compared LLM and fellow performance metrics.
Main Results:
- LLM response consistency varied, with Claude 3.5 Sonnet and Llama 3.2 being most consistent.
- LLMs generally scored higher than fellows (58-85% vs. 57%).
- LLMs showed weak to no correlation with fellow difficulty indices and failed to predict point biserial indices.
Conclusions:
- LLMs have limited utility in predicting specific MCQ psychometric properties.
- Higher-performing LLMs correlated less with human performance, suggesting a potential inverse relationship.
- Future research should focus on LLMs for broader assessment optimization, not item-level prediction.
More Related Videos
03:14Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
09:00Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
Published on: August 16, 2024
Related Concept Videos
Response Surface Methodology
The process of RSM involves several key steps:
Typical Model Studies
Components of Language
Survival Tree
Building a Survival Tree
Constructing a...