Evaluation of Large Language Models in Thailands National Medical Licensing Examination
Saowaprut, P.; Wabina, R. S.; Yang, J.; Siriwat, L.
Show abstract
Advanced general-purpose Large Language Models (LLMs), including OpenAIs Chat Generative Pre-trained Transformer (ChatGPT), Googles Gemini and Anthropics Claude, have demonstrated capabilities in answering clinical questions, including those with image inputs. The Thai National Medical Licensing Examination (ThaiNLE) lacks publicly accessible specialist-confirmed study materials. This study aims to evaluate whether LLMs can accurately answer Step 1 of the ThaiNLE, a test similar to Step 1 of the United States Medical Licensing Examination (USMLE). We utilized a mock examination dataset comprising 300 multiple-choice questions, 10.2% of which included images. LLMs capable of processing both image and text data were used, namely GPT-4, Claude 3 Opus and Gemini 1.0 Pro. Five runs of each model were conducted through their application programming interface (API), with the performance assessed based on mean accuracy. Our findings indicate that all tested models surpassed the passing score, with the top performers achieving scores more than two standard deviations above the national average. Notably, the highest-scoring model achieved an accuracy of 88.9%. The models demonstrated robust performance across all topics, with consistent accuracy in both text-only and image-enhanced questions. However, while the LLMs showed strong proficiency in handling visual information, their performance on text-only questions was slightly superior. This study underscores the potential of LLMs in medical education, particularly in accurately interpreting and responding to a diverse array of exam questions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 97%
- Uncovering the effects of model initialization on deep model generalization: A study with adult and pediatric chest X-ray images 95%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 94%
Similar papers in this journal
- Evaluation of the performance of GPT-3.5 and GPT-4 on the Medical Final Examination 97%
- On evaluation metrics for medical applications of artificial intelligence 93%
- Mitigating Machine Learning Bias Between High Income and Low-Middle Income Countries for Enhanced Model Fairness and Generalizability 92%
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 92%
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 91%
Similar papers in this journal
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 95%
- Large language models for generating medical examinations: systematic review 92%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.