Related Experiment Video
Updated: Apr 21, 2026

Assessment of Viability of Human Fat Injection into Nude Mice with Micro-Computed Tomography
Published on: January 7, 2015
Would AI Take a Research Year? Pilot Study Evaluating the Reliability of ChatGPT in Advising Plastic Surgery
Naomi N Ghahrai1, Skye T Coffey2, Virginia Bailey1
1Department of Plastic Surgery, Maxillofacial, and Oral Health, University of Virginia Health, Charlottesville, VA, USA.
Introduction:
Large language models (LLMs) such as OpenAI's GPT-4o are increasingly used to summarize information and report trends in available data for medical education. For integrated plastic surgery, the utility of LLMs to recommend taking a research year has not been established. We aim to establish the reliability of ChatGPT reproducibility of research year recommendations for medical students applying to integrated plastic surgery.
Methods:
De-identified, self-reported integrated plastics applicant profiles in publicly available Google Sheets from 2022-2025 were assembled. Inputs provided to GPT-4o (three runs per profile) included Step 2 CK (Clinical Knowledge) score, AOA designation, and research productivity. Research-year status and match outcome were withheld. The model returned a binary recommendation to pursue a research year. Reproducibility was summarized as cross-run concordance. We compared model recommendations with applicants' actual research-year decisions.
Results:
Of 98 entries, 55 complete profiles were retained. Mean Step 2 CK was 258.3 (SD = 10.4). Applicants reported a mean 20.1 (SD = 19.9) research presentations, 3.84 (SD = 3.6) first-author publications, and 9.18 (SD = 6.4) total publications. Twenty-one eligible applicants (51.2%) reported AOA. Overall, 98.2% (54/55) matched. Across the three computed runs, there was a 98% concordance in recommendations. The LLM recommended a research year for 32.7% (18/55) of entries, whereas 45.5% (25/55) actually undertook one (p = 0.208). Agreement between model recommendations and applicant decisions was 41.8% (p = 0.28).
Conclusion:
ChatGPT demonstrated internal consistency, but its recommendations could not predict which students would take a research year en route to a successful residency match.
More Related Videos
03:07Single-Port Robotic-assisted Transaxillary Breast-conserving Surgery: A Prospective, Single-arm, Non-randomized Phase IIa Clinical Trial
Published on: August 19, 2025
06:26Quantitative Assessment Protocol for Facial Soft Tissue Volumetric Changes with Stereophotogrammetry
Published on: December 9, 2025