Related Experiment Videos
Artificial Professionalism: An Evaluation of Six Large Language Models on the UK Multi-Specialty Recruitment
Mohammad Jakir Ahmed1, Munsif M Mansoor2, Hamza A Mahmood3
1General Internal Medicine, Watford General Hospital, London, GBR.
Background:
The Professional Dilemmas (PD) paper of the UK Multi-Specialty Recruitment Assessment (MSRA) is a situational judgement test (SJT) used to shortlist candidates for several postgraduate specialty training programmes, including general practice, core psychiatry, clinical radiology and obstetrics and gynaecology. Candidates increasingly use publicly available large language models (LLMs) informally during examination preparation, yet the performance of contemporary LLMs on this examination has not been established. This study aimed to quantify and compare the performance of six publicly available LLMs on the official MSRA practice paper.
Methods:
The 22-item PD section of the NHS England (Workforce, Training and Education) 2023 MSRA official practice paper (maximum 86 marks) was administered in April 2026 to six publicly available LLMs via their free-tier consumer chat interfaces: ChatGPT (OpenAI), Microsoft Copilot, Claude (Anthropic), Gemini (Google), DeepSeek and Grok (xAI). Each item was entered verbatim with no system prompt. Responses were scored against the official answer key. Descriptive statistics, performance by item format, Spearman and Pearson inter-model correlations and a Friedman test for overall between-model differences were computed.
Results:
All six models scored between 51/86 (59.30%) and 61/86 (70.93%), with a mean of 56.8/86 (66.09%; SD 4.06 percentage points). Gemini scored highest at 61/86 (70.93%) and Grok lowest at 51/86 (59.30%), but a Friedman test showed these differences were not statistically significant (χ²(5) = 7.64, p = 0.18). Across all models, mean performance on multiple-choice items at 26.5/33 (80.3%) markedly exceeded performance on ranking items at 30.3/53 (57.2%) (paired t = -16.1, p < 0.001). Models agreed completely on four of 22 items. Similarity in item-level performance patterns between models was modest (Spearman ρ 0.03-0.67).
Conclusions:
Contemporary general-purpose LLMs achieve roughly two-thirds of available marks on the MSRA PD paper with no statistically significant differences detected between them, but reason far less reliably on ranking items than on multiple-choice items. Candidates should not treat any single LLM answer as authoritative and should refer to the primary sources on which the PD paper is built, such as the General Medical Council's Good Medical Practice.
Related Concept Videos
Stereotype Content Model
Halo Effect