Is ChatGPT smarter than Otolaryngology trainees? A comparison study of board style exam questions
Patel, J.; Robinson, P.; Illing, E.; Anthony, B.
Show abstract
ObjectivesThis study compares the performance of the artificial intelligence (AI) platform Chat Generative Pre-Trained Transformer (ChatGPT) to Otolaryngology trainees on board style exam questions. MethodsWe administered a set of 30 Otolaryngology board style questions to medical students (MS) and Otolaryngology residents (OR). 31 MSs and 17 ORs completed the questionnaire. The same test was administered to ChatGPT version 3.5, five times. Comparisons of performance were achieved using a one-way ANOVA with Tukey Post Hoc test, along with a regression analysis to explore the relationship between education level and performance. ResultsThe average scores increased each year from MS1 to PGY5. A one-way ANOVA revealed that ChatGPT outperformed trainee years MS1, MS2, and MS3 (p = <0.001, 0.003, and 0.019, respectively). PGY4 and PGY5 otolaryngology residents outperformed ChatGPT (p = 0.033 and 0.002, respectively). For years MS4, PGY1, PGY2, and PGY3 there was no statistical difference between trainee scores and ChatGPT (p = .104, .996, and 1.000, respectively). ConclusionChatGPT can outperform lower-level medical trainees on Otolaryngology board-style exam but still lacks the ability to outperform higher-level trainees. These questions primarily test rote memorization of medical facts; in contrast, the art of practicing medicine is predicated on the synthesis of complex presentations of disease and multilayered application of knowledge of the healing process. Given that upper-level trainees outperform ChatGPT, it is unlikely that ChatGPT, in its current form will provide significant clinical utility over an Otolaryngologist.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 93%
- Evaluation of Large Language Models in Medical Examinations:A Scoping Review Protocol 93%
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 93%
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 92%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 91%
Similar papers in this journal
- Cross-modal sensory boosting to improve high-frequency hearing loss 90%
- Caregivers’ Perspective: Satisfaction With Healthcare Services At The Paediatric Specialist Clinic Of The National Referral Centre In Malaysia 89%
- Learning from the resilience of hospitals and their staff to the COVID-19 pandemic: a scoping review 89%
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 96%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 95%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.