Related Experiment Video
Updated: Jan 10, 2026

03:14
Augmenting Large Language Models via Vector Embeddings to Improve Domain-Specific Responsiveness
Published on: December 6, 2024
994
Evaluating the performance of large language models on the ASPS In-Service Examination: A comparative analysis with
Ramin Shekouhi1, Mary M Holohan1, Oygul Mirzalieva2
1Division of Plastic and Reconstructive Surgery, Department of Surgery, Louisiana State University Health Sciences Center, New Orleans, LA, USA.
Journal of Plastic, Reconstructive & Aesthetic Surgery : JPRAS
|November 27, 2025
Summary
Leading large language models (LLMs) show high accuracy on the Plastic Surgery In-Service Training Examination (PSITE), often surpassing resident and practitioner performance. This indicates their potential in surgical education.
Area of Science:
- Medical Education
- Artificial Intelligence
- Plastic Surgery
Background:
- Large language models (LLMs) are emerging as potential tools in medical education.
- Their performance in specialized fields like plastic surgery requires thorough evaluation.
Purpose of the Study:
- To assess the accuracy and comparative performance of three leading LLMs: ChatGPT 4.0, DeepSeek V3, and Gemini 2.5.
- To evaluate these LLMs on the American Board of Plastic Surgery Plastic Surgery In-Service Training Examination (PSITE) over a 20-year period.
Main Methods:
- LLMs were tested on PSITE questions spanning two decades.
- Performance was measured by overall accuracy and percentile ranks against normative data of residents and practitioners.
Main Results:
- ChatGPT, DeepSeek, and Gemini demonstrated high and comparable overall accuracy (75.0%, 74.8%, 74.5% respectively), with no statistically significant differences.
- DeepSeek achieved the highest percentile ranks (81st resident, 89th practitioner), closely followed by ChatGPT and Gemini, with no significant differences across LLMs.
Conclusions:
- Modern LLMs exhibit consistent, high-level performance on the PSITE.
- These advanced AI models frequently exceed the median performance of both plastic surgery residents and practitioners, suggesting significant potential in surgical training and education.

