Related Experiment Videos
Estimating the Prevalence of Generative AI Use in Medical School Application Essays: Cross-Sectional Study
Nicholas C Spies1, Valerie S Ratts2, Ian S Hagemann3
1Department of Pathology, University of Utah, Salt Lake City, UT, United States.
Background:
Generative AI tools became widely available to the public in November 2022. The extent to which these tools have been used by medical school applicants during the admissions process is unknown.
Objective:
We aimed to estimate the extent of generative AI use among cohorts of applicants spanning the rollout of these tools.
Methods:
We retrospectively analyzed 6000 essays from 2364 applicants submitted to a US medical school in 2021 to 2022 (baseline, before the wide availability of AI) and 2023 to 2024 (test year) to estimate the prevalence of AI use and its relation to other application data. We used GPTZero, a commercially available detection tool, to generate a metric (Phuman) reflecting the predicted probability that each essay was completely human generated, ranging from 0 (the essay appears to be entirely AI generated) to 1 (the essay appears to be entirely human generated).
Results:
Fully human-generated negative controls demonstrated a median Phuman of 0.93 (range 0.89-0.97), while fully AI-generated positive controls demonstrated a median Phuman of 0.01 (range 0.00-0.01). The "Personal Comments" essays submitted in the 2023 to 2024 application cycle had a median Phuman of 0.77 (95% CI 0.76-0.78) compared with 0.83 (95% CI 0.82-0.85) during the 2021 to 2022 cycle. Approximately 12.3% and 2.7% of essays were evaluated as having Phuman <0.5 in the test and baseline years, respectively. Essays submitted as part of the secondary application demonstrated lower Phuman values than those of the American Medical College Application Service (AMCAS) "Personal Comments" essays. In applicant-clustered, multivariable generalized estimating equation analyses, supplementary essay type and younger age were significantly associated with lower Phuman. Application completion date, self-reported gender, program type (MD vs MD-PhD), grade point average (GPA), Medical College Admission Test (MCAT) score, socioeconomic status, and undergraduate major were not significant predictors after false discovery rate correction. Phuman was not predictive of interview invitation or acceptance in adjusted applicant-level logistic regression analyses.
Conclusions:
An AI detection algorithm identified signs of increased use of generative AI in 2023 to 2024 medical school admission applications compared to those in the 2021 to 2022 baseline period, before AI was widely available. AI use did not appear to confer an admissions advantage. Although these results provide information about the applicant pool as a whole, AI detection is imperfect. We do not recommend deploying AI detection for individual applications in live admissions cycles.