Back

ChatGPT and the Clinical Informatics Board Examination: The End of Knowledge-based Medical Board Maintenance?

Kumah-Crystal, Y.; Mankowitz, S.; Embi, P.; Lehmann, C. U.

2023-04-28 health informatics
10.1101/2023.04.25.23289105 medRxiv
Show abstract

ObjectivesAssess ChatGPTs performance on the Clinical Informatics Board Examination (CIBE) and discuss the implications of large language models (LLMs) for board certification and maintenance. Materials and MethodsWe tested ChatGPT using 260 multiple-choice questions from Mankowitzs Clinical Informatics Board Review book, omitting six image-dependent questions. ResultsChatGPT answered 190 (74%) of 254 eligible questions correctly. While performance varied across the Clinical Informatics Core Content Areas, differences were not statistically significant. DiscussionChatGPTs performance raises concerns about the potential misuse in medical certification and the future validity of knowledge assessment exams. While ChatGPT is able to answer multiple-choice questions accurately, relying on AI systems for exams will compromise the credibility and validity of at-home assessments and undermine public trust. ConclusionThe advent of AI and LLMs threatens to upend existing processes to board certification and maintenance and necessitates new approaches to the evaluation of proficiency in medical education.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.