Related Experiment Video
Updated: Sep 13, 2026

Reliability of Artificial Intelligence-Based Cone Beam Computed Tomography Integration with Digital Dental Images
Published on: February 23, 2024
Human Judgment and the Limits of Artificial Intelligence for Automated Rank Order Lists in Diagnostic Radiology
Cody H Savage1, Rydhwana Hossain2, James Mac Tonascia3
1Department of Diagnostic Radiology and Nuclear Medicine, University of Maryland Medical Center, Baltimore, MD, USA.
Objective:
To evaluate whether large language models (LLMs) can reliably reproduce a residency program's rank order list (ROL) from pre-interview application data alone versus with human interviews included, and to determine whether automated ranking tools safely mirror human consensus or inadvertently introduce systematic displacements against specific applicant subgroups.
Methods:
In this single-institution pilot study of 148 applicants interviewed during the 2025-2026 cycle, seven LLM configurations (three proprietary models [each at "medium" and "high" reasoning levels], and one open-weight model) each generated 10 independent ROLs from de-identified application data under two conditions: pre-interview application-only data alone, and then with interviewer scores added (140 lists total). Non-model baselines (sorting solely by interview score or USMLE Step 2 CK score) were also established. Lists were compared with the committee's final ROL using a truth-anchored weighted Kendall's τ and precision@k (k=10-40); rank displacement by applicant subgroup was also assessed.
Results:
Producing the committee's ROL required approximately 390 faculty-hours to order 148 applicants for 7 positions. In the application-only condition, all LLM configurations showed low agreement with the final ROL (median τ 0.15-0.36). In the positive control condition with human interviews included, agreement rose to excellent (median τ 0.84-0.93). Sorting applicants by summed interviewer score alone reproduced the final list at τ-b=0.83, whereas sorting by USMLE Step 2 CK score alone produced τ-b=0.12. Increasing reasoning level from "medium" to "high" did not consistently improve agreement. In the application-only condition, LLM rankings systematically placed female, international, and non-MD applicants below the human committee's list (q<0.05), while program signalers were placed higher.
Discussion:
LLMs failed to replicate the program's final ROL from pre-interview application-only data alone. The positive control including the human interview confirms this human-LLM lack of agreement may stem from the limited application-only data provided, rather than the models' inability to perform ranking tasks once provided the additional human judgment information. Importantly, these findings do not establish whether the LLM model or the human committee produced the "better" rank order list; this data suggests that LLM rankings generated solely from pre-interview data systematically differed from the human committee's consensus, and may disproportionately penalize certain subgroups. Our findings suggest that AI should not be used as a primary ROL generator or a replacement for human judgment, however may serve supportive roles in tasks such as data retrieval and retrospective auditing.
