Related Experiment Video
Updated: May 28, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Artificial Intelligence in Medical Assessment: Reliability and Performance of Multimodal Large Language Models in a
Ibrahim Güler1,2,3, Gerrit Grieb2,4, Armin Kraus1
1Department of Plastic, Aesthetic and Hand Surgery, Otto-von-Guericke University, 39120 Magdeburg, Germany.
Behavioral Sciences (Basel, Switzerland)
|May 27, 2026
Summary
Large language models (LLMs) show high accuracy and reproducibility in medical licensing exams. These AI tools demonstrate reliable performance, supporting their use in assessment research.
Area of Science:
- Medical Education
- Artificial Intelligence
- Psychometrics
Background:
- Artificial intelligence (AI) integration into assessments is growing.
- Limited evidence exists on large language models' (LLMs) reliability in high-stakes evaluations.
Purpose of the Study:
- To evaluate the performance and reproducibility of multimodal LLMs in a medical assessment.
- To analyze measurement properties of LLMs in a standardized testing environment.
Main Methods:
- A cross-sectional dual-setup design used a national medical licensing examination (240 items).
- Ten LLMs were tested in Setup 1 (single run); six were tested in Setup 2 (five runs each).
- Analyzed accuracy, inter-run agreement (Cohen's kappa), and paired comparisons (McNemar's test).
Main Results:
- LLM accuracy ranged from 72.08% to 92.92%.
- All models exhibited near-perfect inter-run agreement (mean kappa ≥ 0.96) with low variability.
- Performance on image-based items was comparable to or higher than text-only items.
Conclusions:
- Multimodal LLMs achieve high accuracy and reproducibility in large-scale assessments.
- Findings support LLMs as subjects for AI-based assessment research.
- Cognitive equivalence to human performance requires further investigation beyond accuracy metrics.