Back

Evaluating ChatGPT4 in Canadian Otolaryngology-Head and Neck Surgery Board Examination using the CVSA Model

Long, C.; Lowe, K.; dos Santos, A.; Zhang, J.; Alanazi, A.; O'Brien, D.; Wright, E.; Cote, D.

2023-06-01 otolaryngology
10.1101/2023.05.30.23290758 medRxiv
Show abstract

BackgroundChatGPT is among the most popular Large Language Models (LLM), exhibiting proficiency in various standardized tests, including multiple-choice medical board examinations. However, its performance on Otolaryngology-Head and Neck Surgery (OHNS) board exams and open-ended medical board examinations has not been reported. We present the first evaluation of LLM (ChatGPT-4) on such examinations and propose a novel method to assess an artificial intelligence (AI) models performance on open-ended medical board examination questions. MethodsTwenty-one open end questions were adopted from the Royal College of Physicians and Surgeons of Canadas sample exam to query ChatGPT-4 on April 11th, 2023, with and without prompts. A new CVSA (concordance, validity, safety, and accuracy) model was developed to evaluate its performance. ResultsIn an open-ended question assessment, ChatGPT-4 achieved a passing mark (an average of 75% across three trials) in the attempts. The model demonstrated high concordance (92.06%) and satisfactory validity. While demonstrating considerable consistency in regenerating answers, it often provided only partially correct responses. Notably, concerning features such as hallucinations and self-conflicting answers were observed. ConclusionChatGPT-4 achieved a passing score in the sample exam, and demonstrated the potential to pass the Canadian Otolaryngology-Head and Neck Surgery Royal College board examination. Some concerns remain due to its hallucinations that could pose risks to patient safety. Further adjustments are necessary to yield safer and more accurate answers for clinical implementation.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.