Related Experiment Video
Updated: Aug 29, 2026

Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
Large Language Models as Simulated Candidates in Objective Structured Clinical Examinations: A Rubric-Mapping
Elise Lupon1,2,3, Alexandre O Gérard1,4, Alexandre Destere1,4,5
1Département d'Innovation Pédagogique et Intelligence Artificielle, Université Côte d'Azur, Nice, France.
Background:
Large language models (LLMs) perform well on knowledge-based medical examinations, but evidence on their utility for performance-based assessments such as Objective Structured Clinical Examinations (OSCEs) remains limited. Beyond simulating candidate answers, LLMs could support quality assurance of OSCE stations by auditing the completeness and internal consistency of analytic scoring rubrics.
Methods:
In this exploratory, single-run study, GPT-4 Turbo (OpenAI, accessed via the ChatGPT Plus consumer interface, May 2025) simulated a final-year medical student (DFASM3) across 10 validated French national OSCE stations covering all 11 nationally defined competency domains, using a fixed system prompt applied identically across stations. Each station was run once in a baseline condition and once in a reference-augmented (RAG) condition, in which official LiSA learning objectives and Starting Clinical Situation (SCS) documents were additionally provided. Outputs were scored against official analytic rubrics by two clinical educators through independent review (Cohen's κ=0.84), with disagreements resolved by consensus; off-checklist content was similarly adjudicated.
Results:
Mean baseline checklist coverage was 81% (SD 10.2), with A-level item omission averaging 10% (SD 7.2). RAG coverage increased to 95% (SD 5.1; Wilcoxon signed-rank test, p=0.002), recovering 1-4 previously missed items per station. Off-checklist additions averaged 22% of checklist items per station, nearly all clinically relevant; one incorrect addition was identified. Coverage was higher in structured, action-oriented stations (91.6% ± 4.1) than dialogue-heavy stations (73.7% ± 7.4). Given the single-run design and absence of a human comparator, results are exploratory and hypothesis-generating rather than robust performance estimates.
Conclusion:
This proof-of-concept study shows that an LLM can generate structured, rubric-mappable OSCE responses and may help educators flag missing, ambiguous, or inconsistent checklist elements for expert review. These exploratory findings support further investigation of LLMs as decision-support tools for OSCE station pre-validation, pending replication with multiple runs, larger station banks, and human comparators.