ChatGPT and the Clinical Informatics Board Examination: The End of Knowledge-based Medical Board Maintenance?
Kumah-Crystal, Y.; Mankowitz, S.; Embi, P.; Lehmann, C. U.
Show abstract
ObjectivesAssess ChatGPTs performance on the Clinical Informatics Board Examination (CIBE) and discuss the implications of large language models (LLMs) for board certification and maintenance. Materials and MethodsWe tested ChatGPT using 260 multiple-choice questions from Mankowitzs Clinical Informatics Board Review book, omitting six image-dependent questions. ResultsChatGPT answered 190 (74%) of 254 eligible questions correctly. While performance varied across the Clinical Informatics Core Content Areas, differences were not statistically significant. DiscussionChatGPTs performance raises concerns about the potential misuse in medical certification and the future validity of knowledge assessment exams. While ChatGPT is able to answer multiple-choice questions accurately, relying on AI systems for exams will compromise the credibility and validity of at-home assessments and undermine public trust. ConclusionThe advent of AI and LLMs threatens to upend existing processes to board certification and maintenance and necessitates new approaches to the evaluation of proficiency in medical education.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 95%
- A national professional development program fills mentoring gaps for postdoctoral researchers 93%
- Assessment of research ethics education offerings of pharmacy master programs: a qualitative content analysis 93%
Similar papers in this journal
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 93%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 92%
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 96%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 95%
- Evaluation of Statistical Illiteracy in Latin American Clinicians and of the Efficacy of a 10-Hour Course 94%
Similar papers in this journal
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 93%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.