Estimating the Prevalence of Generative AI Use in Medical School Application Essays
Hagemann, I. S.; Ratts, V. S.; Spies, N. C.
Show abstract
BackgroundGenerative artificial intelligence (AI) tools became widely available to the public in November 2022. The extent to which these tools are being used by aspiring medical school applicants during the admissions process is unknown. MethodsWe retrospectively analyzed 6,000 essays submitted to a U.S. medical school in 2021- 2022 (baseline, before wide availability of AI) and in 2023-2024 (test year) to estimate the prevalence of AI use and its relation to other application data. We used GPTZero, a commercially available detection tool, to generate a metric for the likelihood that each essay was human-generated, Phuman, ranging from 0 (entirely AI) to 1 (entirely human). ResultsFully human-generated negative controls demonstrated a median Phuman of 0.93, while AI-generated positive controls demonstrated a median Phuman of 0.01. Personal Comments essays submitted in the 23- 24 cycle had a median human-generated score of 0.77 (95% confidence interval 0.76-0.78), versus 0.83 (95% CI 0.82-0.85) during the 21- 22 cycle. Approximately 12.3 and 2.7% of essays were evaluated as having Phuman < 0.5 in the test and baseline year, respectively. Secondary essays demonstrated lower Phuman than Personal Comments essays, suggesting more AI use. In multivariate analysis, younger age, visa requirement, and higher GPA were significantly associated with lower Phuman. No differences were observed in gender, MCAT score, undergraduate major, or socioeconomic status. Phuman was not predictive of admissions outcomes in uni- or multivariate analyses. ConclusionsAn AI detection algorithm estimated significantly increased use of generative AI in 2023-2024 medical school admission applications, as compared to the 2021-2022 baseline. Estimated AI use demonstrated no significant differences in admissions decisions. While these results provide information about the applicant pool as a whole, AI detection is imperfect. We recommended exercising caution before deploying any AI detection tools on individual applications in live admissions cycles. DescriptionMedical school applicants increased their use of generative AI to write application essays in the most recent admissions cycle, but this use did not confer an admissions advantage.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 92%
- Inferring Gender from First Names: Comparing the Accuracy of Genderize, Gender API, and the gender R Package on Authors of Diverse Nationality 91%
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 92%
- Fitness tracking reveals task-specific associations between memory, mental health, and physical activity 90%
- Open-loop lab-on-a-chip technology enables remote computer science training in Latinx life sciences students 90%
Similar papers in this journal
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 91%
- Anticipatory Emotions and Academic Performance: The Role of Boredom in a Preservice Teachers' Lab Experience 88%
- Linear vector models of time perception account for saccade and stimulus novelty interactions 88%
Similar papers in this journal
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 91%
- Large language models for generating medical examinations: systematic review 90%
- Making the Match or Breaking it? Values, Perceptions, and Obstacles of Trainees Applying into Physician-Scientist Training Programs 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.