Related Experiment Video
Updated: Jun 4, 2025

An Experimental Analysis of Children's Ability to Provide a False Report about a Crime
Published on: May 3, 2016
Can Artificial Intelligence Deceive Residency Committees? A Randomized Multicenter Analysis of Letters of
Samuel K Simister1, Eric G Huish, Eugene Y Tsai
1From the University of California, Davis, Sacramento, CA (Simister, Le, Meehan, Leshikar, Saiz, and Lum), the San Joaquin General Hospital, French Camp, CA (Huish), the Cedars Sinai, Los Angeles, CA (Tsai), and the Yale University, New Haven, CT (Halim and Tuason).
Generative artificial intelligence (AI) can produce letters of recommendation (LORs) indistinguishable from human authors. Residency selection committees struggle to differentiate AI-generated LORs, impacting application evaluations.
Area of Science:
- Medical Education
- Artificial Intelligence in Medicine
- Surgical Residency Admissions
Background:
- Generative artificial intelligence (AI) presents novel challenges and opportunities in medical education, particularly for residency applications.
- The increasing sophistication of AI necessitates an evaluation of its impact on traditional application components like letters of recommendation (LORs).
Purpose of the Study:
- To assess the capability of orthopaedic surgery residency selection committee members to discern between human-written and AI-generated letters of recommendation (LORs).
- To evaluate the accuracy of identifying AI-generated LORs from various platforms (ChatGPT, Google BARD) compared to human authors.
Main Methods:
- A multicenter, single-blind trial involving 45 LORs (15 human, 15 ChatGPT, 15 Google BARD).
- Seven faculty reviewers from four residency programs evaluated LORs based on 11 standardized characteristics and ranked applicants on a 1-10 scale.
- Statistical analysis included ordinal regression and receiver operating characteristic (ROC) curves to determine accuracy in identifying AI authorship.
Main Results:
- Faculty reviewers incorrectly identified AI-generated letters 63% of the time and human-generated letters 40% of the time.
- No significant increase in identification accuracy was observed over time.
- Google BARD demonstrated superior performance compared to human authors in accuracy, adaptability, and perceived commitment.
Conclusions:
- Faculty members were unable to reliably distinguish between human and AI-generated LORs, indicating AI's capacity to produce comparable recommendation letters.
- The findings underscore the need for residency selection committees to re-evaluate the role and weight of LORs in the application process.
- Further research is warranted to develop robust methods for verifying LOR authenticity in the era of AI.
More Related Videos
06:05Employing Transcranial Magnetic Stimulation in a Resource Limited Environment to Establish Brain-Behavior Relationships
Published on: April 20, 2022
07:10Depletion of Mouse Cells from Human Tumor Xenografts Significantly Improves Downstream Analysis of Target Cells
Published on: July 29, 2016
Related Concept Videos
Ethics in Research
False Memories
One primary source of false memories is misattribution, where individuals incorrectly associate external information...
The Representativeness Heuristic
Stereotype Threat and Self-fulfilling Prophecies
Ethical Dilemmas II
Stereotype Content Model