Related Experiment Video
Updated: Oct 1, 2026

Introduction of an Integrated Pathology Image Management, Artificial Intelligence, and Reporting System
Published on: July 11, 2025
The Use of AI Software in Reviewing Applicants for General Surgery Residency
Kylie Dickerson1, Logan Kenny1, Jennifer Preston1
1Department of Surgery, University of Arizona College of Medicine Phoenix, Phoenix, Arizona.
Background:
Artificial Intelligence (AI) tools are increasingly used to improve efficiency, including in residency application review. While these technologies may streamline processes, their ability to identify optimal candidates for specific programs remains uncertain. This study aims to evaluate if AI can enhance the productivity of residency application review without compromising selection quality.
Methods:
General surgery residency applications were analyzed using a large language model-based AI platform configured with customized program-specific weighting of applicant metrics. The AI-generated scores and rank list were compared with results from the traditional review process in which the application review committee manually reviews applications, scores applicants, and selects candidates for interview.
Results:
A total of 847 applications were reviewed, and 96 applicants were manually selected (MS) for interviews by the application review committee. Of these, 32 (33%) were also ranked in the AI's top 96 candidates, including 7 in the top 10. Quartile analysis demonstrated that 51 (53.1%) MS candidates were in the top quartile of AI rankings, 26 (27.1%) in the second, 8 (8.3%) in the third, and 11 (11.4%) in the bottom. Adjusting the weights of various factors did not significantly alter these distributions. Technical review revealed scoring anomalies, including negligible AI scores for some candidates selected for interview. Analysis of the weighted factors showed that letters of recommendation contributed 29.4% of MS vs 25.1% of AI scoring (p = 0.009), research 6.3% vs 8.0% (p = 0.017), MSPE 7.4% vs 9.0% (p = 0.073), and factors such as USMLE and experiences were equivalent between the 2 groups.
Conclusion:
AI-assisted review may serve as a useful adjunct for initial screening but requires further refinement and training to better replicate human qualitative judgment. Future work will include evaluating concordance between AI rankings and the final program rank list and additional application review to guide software optimization and improve reliability as a screening tool.