Related Experiment Video
Updated: Mar 24, 2026

Author Spotlight: Implementing the Enhanced Recovery After Surgery Concept in Rehabilitation Following Anterior Cruciate Ligament Reconstruction
Published on: March 1, 2024
Does ChatGPT Show Gender Bias When Drafting Letters of Recommendation for Applicants to Orthopaedic Surgery Residency
Eric Mao1, Eve R Glenn1, Dawn M LaPorte1
1Department of Orthopaedic Surgery, The Johns Hopkins University, Baltimore, Maryland.
Introduction:
Implicit gender biases in letters of recommendation (LoRs) may differentially influence the success of applicants to residency positions. With the inevitable use of large language models (LLMs) such as ChatGPT to draft LoRs, concerns have emerged regarding whether these models reproduce human biases in their writing. The purpose of this study was to examine whether ChatGPT exhibits gender bias when drafting LoRs for hypothetical orthopaedic surgery residency applicants.
Methods:
Thirty paired prompts were created describing a variety of mentor-mentee relationships, manipulating only the applicant's gendered name and pronouns while holding all other factors constant. Prompts were sequentially input into ChatGPT-5.0, and output LoRs were saved. Linguistic Inquiry and Word Count (LIWC) software was used to characterize LoRs across 4 summary measures and 28 word categories. Paired t tests were used to compare the composition of male and female letters across these dimensions. Word counts were compared similarly.
Results:
The mean length of recommendations for men was 304 ± 53 words. For women, the mean length was 310 ± 47 words. There was no significant difference in word count between groups (p = 0.364). However, recommendations for women were composed of more auxiliary verbs (4.67% ± 1.10% vs. 4.17% ± 1.02%; p = 0.045), communication-related words (0.80% ± 0.51% vs. 0.59% ± 0.46%; p = 0.047), and personal pronouns (10.48% ± 1.18% vs. 9.82% ± 0.87%; p = 0.005) than recommendations for men. A follow-up analysis using a gender-neutral name, while only varying pronouns between prompts demonstrated that recommendations for women were composed of more "prosocial" words than recommendations for men (3.27% ± 1.16% vs. 2.79% ± 1.00%; p = 0.003).
Conclusion:
ChatGPT-assisted drafting of LoRs includes nuanced and systematic gender-based linguistic differences. For orthopaedic letter writers, the use of LLMs must be accompanied by structured review, bias-aware training, and standardized templates to avoid inadvertently perpetuating inequities.
More Related Videos
09:51The Transition to an Anterior-Based Muscle Sparing Approach Improves Early Postoperative Function but is Associated with a Learning Curve
Published on: September 7, 2022
08:15Three-Dimensional Preoperative Virtual Planning in Derotational Proximal Femoral Osteotomy
Published on: February 17, 2023
Related Concept Videos
Confirmation Biases
Stereotypes, Prejudice, and Discrimination