Related Experiment Video
Updated: Jun 9, 2025

One Dimensional Turing-Like Handshake Test for Motor Intelligence
Published on: December 15, 2010
Can Artificial Intelligence Fool Residency Selection Committees? Analysis of Personal Statements by Real Applicants
Zachary C Lum1,2, Lohitha Guntupalli1, Augustine M Saiz2
1Department of Surgery, Kiran Patel School of Osteopathic and Allopathic Medicine, Nova Southeastern University, Davie, Florida.
Generative artificial intelligence (AI) tools like ChatGPT and Google BARD can create convincing personal statements for residency applications. However, faculty reviewers could distinguish real statements from AI-generated ones, with ChatGPT being more deceptive.
Area of Science:
- Medical Education
- Artificial Intelligence in Healthcare
- Residency Admissions
Background:
- Generative artificial intelligence (AI) capabilities are rapidly advancing.
- The use of AI in crafting personal statements for medical residency applications is an emerging area.
- Understanding the detectability and quality of AI-generated personal statements is crucial for admissions committees.
Purpose of the Study:
- To evaluate the ability of generative AI tools (ChatGPT and Google BARD) to create personal statements for medical residency applications.
- To assess whether faculty on residency selection committees can differentiate between real and AI-generated statements.
- To determine if specific metrics of personal statements are identifiable as AI-generated or real.
Main Methods:
- Generated 15 unique personal statements using ChatGPT and 15 using Google BARD, based on 15 real statements.
- Presented a total of 45 randomized and blinded statements (15 real, 15 ChatGPT, 15 BARD) to faculty reviewers.
- Faculty assessed statements on 14 metrics, including identifying them as real or AI-generated.
Main Results:
- Faculty correctly identified 88% of real statements and 90% of BARD-generated statements, but only 44% of ChatGPT-generated statements.
- Overall accuracy in identifying AI vs. real statements was 89% for BARD and 74% when including ChatGPT.
- BARD-generated statements performed significantly poorer than real and ChatGPT statements across all metrics (p < 0.001).
- Real statements were superior to ChatGPT statements in Personal Interests, Reasons for Choosing Residency, Career Goals, Compelling Nature, and Originality (p < 0.001).
Conclusions:
- Faculty can accurately identify real and BARD-generated personal statements, but ChatGPT successfully deceived reviewers 56% of the time.
- While AI can produce convincing statements, replicating humanistic elements like personal nuances and individual experiences remains challenging.
- Residency selection committees should consider emphasizing metrics reflecting genuine personal experience, such as interests, motivations, and originality, when evaluating applications.
Related Concept Videos
Reliability and Validity
Stereotype Threat and Self-fulfilling Prophecies
The Representativeness Heuristic
Cause and Effect
Hypothesis: Accept or Fail to Reject?
There are two ways to indicate that the null hypothesis is not rejected. 'Accept' the null...
Confirmation Biases

