How does ChatGPT4 preform on Non-English National Medical Licensing Examination? An Evaluation in Chinese Language
Fang, C.; Ling, J.; Zhou, J.; Wang, Y.; Liu, X.; Jiang, Y.; Wu, Y.; Chen, Y.; Zhu, Z.; Ma, J.; Yan, Z.; Yu, P.; Liu, X.
Show abstract
BackgroundChatGPT, an artificial intelligence (AI) system powered by large-scale language models, has garnered significant interest in the healthcare. Its performance dependent on the quality and amount of training data available for specific language. This study aims to assess the of ChatGPTs ability in medical education and clinical decision-making within the Chinese context. MethodsWe utilized a dataset from the Chinese National Medical Licensing Examination (NMLE) to assess ChatGPT-4s proficiency in medical knowledge within the Chinese language. Performance indicators, including score, accuracy, and concordance (confirmation of answers through explanation), were employed to evaluate ChatGPTs effectiveness in both original and encoded medical questions. Additionally, we translated the original Chinese questions into English to explore potential avenues for improvement. ResultsChatGPT scored 442/600 for original questions in Chinese, surpassing the passing threshold of 360/600. However, ChatGPT demonstrated reduced accuracy in addressing open-ended questions, with an overall accuracy rate of 47.7%. Despite this, ChatGPT displayed commendable consistency, achieving a 75% concordance rate across all case analysis questions. Moreover, translating Chinese case analysis questions into English yielded only marginal improvements in ChatGPTs performance (P =0.728). ConclusionChatGPT exhibits remarkable precision and reliability when handling the NMLE in Chinese language. Translation of NMLE questions from Chinese to English does not yield an improvement in ChatGPTs performance.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
Similar papers in this journal
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 97%
- Large language models for generating medical examinations: systematic review 95%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 93%
Similar papers in this journal
Similar papers in this journal
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 91%
- Perceptions of Complementary, Alternative, and Integrative Medicine: Insights from a Large-Scale International Cross-Sectional Survey of Surgery Researchers and Clinicians 90%
- Linear vector models of time perception account for saccade and stimulus novelty interactions 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.