Evaluating ChatGPT4 in Canadian Otolaryngology-Head and Neck Surgery Board Examination using the CVSA Model
Long, C.; Lowe, K.; dos Santos, A.; Zhang, J.; Alanazi, A.; O'Brien, D.; Wright, E.; Cote, D.
Show abstract
BackgroundChatGPT is among the most popular Large Language Models (LLM), exhibiting proficiency in various standardized tests, including multiple-choice medical board examinations. However, its performance on Otolaryngology-Head and Neck Surgery (OHNS) board exams and open-ended medical board examinations has not been reported. We present the first evaluation of LLM (ChatGPT-4) on such examinations and propose a novel method to assess an artificial intelligence (AI) models performance on open-ended medical board examination questions. MethodsTwenty-one open end questions were adopted from the Royal College of Physicians and Surgeons of Canadas sample exam to query ChatGPT-4 on April 11th, 2023, with and without prompts. A new CVSA (concordance, validity, safety, and accuracy) model was developed to evaluate its performance. ResultsIn an open-ended question assessment, ChatGPT-4 achieved a passing mark (an average of 75% across three trials) in the attempts. The model demonstrated high concordance (92.06%) and satisfactory validity. While demonstrating considerable consistency in regenerating answers, it often provided only partially correct responses. Notably, concerning features such as hallucinations and self-conflicting answers were observed. ConclusionChatGPT-4 achieved a passing score in the sample exam, and demonstrated the potential to pass the Canadian Otolaryngology-Head and Neck Surgery Royal College board examination. Some concerns remain due to its hallucinations that could pose risks to patient safety. Further adjustments are necessary to yield safer and more accurate answers for clinical implementation.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 92%
Similar papers in this journal
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 94%
- Evaluating user experience with immersive technology in simulation-based education: a modified Delphi study with qualitative analysis 94%
Similar papers in this journal
Similar papers in this journal
- Learning from the resilience of hospitals and their staff to the COVID-19 pandemic: a scoping review 90%
- Caregivers’ Perspective: Satisfaction With Healthcare Services At The Paediatric Specialist Clinic Of The National Referral Centre In Malaysia 90%
- Challenges in Implementing a Mobile AI Chatbot Intervention for Depression Among Youth on Psychiatric Waiting Lists: A Randomized Control Study Termination Report 90%
Similar papers in this journal
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 93%
- Extraction of Crohn's Disease Clinical Phenotypes from Clinical Text Using Natural Language Processing 91%
- Cancer risk algorithms in primary care: can they improve risk estimates and referral decisions? 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.