Jove
Visualize
Contact Us
JoVE
x logofacebook logolinkedin logoyoutube logo
ABOUT JoVE
OverviewLeadershipBlogJoVE Help Center
AUTHORS
Publishing ProcessEditorial BoardScope & PoliciesPeer ReviewFAQSubmit
LIBRARIANS
TestimonialsSubscriptionsAccessResourcesLibrary Advisory BoardFAQ
RESEARCH
JoVE JournalMethods CollectionsJoVE Encyclopedia of ExperimentsArchive
EDUCATION
JoVE CoreJoVE BusinessJoVE Science EducationJoVE Lab ManualFaculty Resource CenterFaculty Site
Terms & Conditions of Use
Privacy Policy
Policies

Related Concept Videos

Response Surface Methodology01:16

Response Surface Methodology

904
Response Surface Methodology (RSM) is a collection of statistical and mathematical techniques used to develop, improve, and optimize processes. It is particularly valuable when many input variables or factors potentially influence a response variable.
The process of RSM involves several key steps:
904
Typical Model Studies01:30

Typical Model Studies

842
Fluid mechanics model studies often utilize scaled-down systems to predict fluid behavior in full-scale environments, such as river flows, dam spillways, and structures interacting with open surfaces. Maintaining Froude number similarity in river models is crucial, as it replicates surface flow features like wave patterns and velocities.
842
Components of Language01:24

Components of Language

840
Language, whether spoken, signed, or written, consists of specific components: lexicon and grammar. The lexicon is the vocabulary of a language, comprising its words. Grammar is the set of rules used to convey meaning through the lexicon. For example, English grammar adds “-ed” to most verbs to indicate past tense. Words are formed by combining phonemes, which are the basic sound units of a language. Different languages have different sets of phonemes (e.g., “ah” vs.
840
Survival Tree01:19

Survival Tree

499
Survival trees are a non-parametric method used in survival analysis to model the relationship between a set of covariates and the time until an event of interest occurs, often referred to as the "time-to-event" or "survival time." This method is particularly useful when dealing with censored data, where the event has not occurred for some individuals by the end of the study period, or when the exact time of the event is unknown.
 Building a Survival Tree
Constructing a...
499

You might also read

Related Articles

Articles linked to this work by shared authors, journal, and citation graph.

Sort by
Same author

Longitudinal 'oral board exams' to maintain educational continuity.

Medical education·2024
Same author

Propofol-Based Anesthesia Maintenance and/or Volatile Anesthetics during Intracranial Aneurysm Repair: A Comparative Analysis of Neurological Outcomes.

Journal of clinical medicine·2023
Same author

Perception of Web-Based Didactic Activities During the COVID-19 Pandemic Among Anesthesia Residents: Pilot Questionnaire Study.

JMIR medical education·2022
See all related articles

Related Experiment Video

Updated: May 3, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.0K

The answer may vary: large language model response patterns challenge their use in test item analysis.

Lauren K Buhl1

  • 1Department of Anesthesiology, Dartmouth Hitchcock Medical Center, Lebanon, NH, USA.

Medical Teacher
|May 4, 2025
PubMed
Summary

Large language models (LLMs) show limited ability to predict multiple-choice question (MCQ) performance metrics like difficulty and point biserial indices. Consistency of LLM responses is key for assessment development, not prediction of item characteristics.

Keywords:
Item analysisartificial intelligence in educationmedical education assessmentresponse variabilitytest development

More Related Videos

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

461
Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

614

Related Experiment Videos

Last Updated: May 3, 2026

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment
06:48

Lexical Decision Task for Studying Written Word Recognition in Adults with and without Dementia or Mild Cognitive Impairment

Published on: June 25, 2019

9.0K
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
03:14

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness

Published on: December 6, 2024

461
Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education
09:00

Author Spotlight: Validation of SICOLE-R for Assessing Cognitive and Reading Skills in Spanish-Speaking Children and Its Role in Personalized Education

Published on: August 16, 2024

614

Area of Science:

  • Medical Education
  • Artificial Intelligence in Assessment

Background:

  • Validating multiple-choice questions (MCQs) requires extensive testing.
  • Large language models (LLMs) offer potential for streamlining assessment development.
  • Predicting psychometric properties of MCQs is a key challenge.

Purpose of the Study:

  • To investigate LLMs' ability to predict MCQ difficulty and point biserial indices.
  • To assess if LLMs can reduce the need for preliminary test population analysis.
  • To compare LLM performance with human expert assessment.

Main Methods:

  • Sixty anesthesiology MCQs were administered to five LLMs and clinical fellows.
  • LLM response patterns, difficulty indices, and point biserial indices were analyzed.
  • Spearman correlation coefficients compared LLM and fellow performance metrics.

Main Results:

  • LLM response consistency varied, with Claude 3.5 Sonnet and Llama 3.2 being most consistent.
  • LLMs generally scored higher than fellows (58-85% vs. 57%).
  • LLMs showed weak to no correlation with fellow difficulty indices and failed to predict point biserial indices.

Conclusions:

  • LLMs have limited utility in predicting specific MCQ psychometric properties.
  • Higher-performing LLMs correlated less with human performance, suggesting a potential inverse relationship.
  • Future research should focus on LLMs for broader assessment optimization, not item-level prediction.