A comparative analysis of the performance of chatGPT4, Gemini Gemini and Claude Claude for the Polish Medical Final Diploma Exam and Medical-Dental Verification Exam.
Wojcik, D.; Adamiak, O.; Czerepak, G.; Tokarczuk, O.; Szalewski, L.
Show abstract
In the realm of medical education, the utility of chatbots is being explored with growing interest. One pertinent area of investigation is the performance of these models on standardized medical examinations, which are crucial for certifying the knowledge and readiness of healthcare professionals. In Poland, dental and medical students have to pass crucial exams known as LDEK (Medical-Dental Final Examination) and LEK (Medical Final Examination) exams respectively. The primary objective of this study was to conduct a comparative analysis of chatbots: ChatGPT-4, Gemini and Claude to evaluate their accuracy in answering exam questions of the LDEK and the Medical-Dental Verification Examination (LDEW), using queries in both English and Polish. The analysis of Model 2, which compared chatbots within question groups, showed that the chatbot Claude achieved the highest probability of accuracy for all question groups except the area of prosthetic dentistry compared to ChatGPT-4 and Gemini. In addition, the probability of a correct answer to questions in the field of integrated medicine is higher than in the field of dentistry for all chatbots in both prompt languages. Our results demonstrate that Claude achieved the highest accuracy in all areas analysed and outperformed other chatbots. This suggests that Claude has significant potential to support the medical education of dental students. This study showed that the performance of chatbots varied depending on the prompt language and the specific field. This highlights the importance of considering language and specialty when selecting a chatbot for educational purposes.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 94%
- Burnout: A Predictor Of Oral Health Impact Profile Among Nigerian Early Career Doctors 94%
- Assessment of research ethics education offerings of pharmacy master programs: a qualitative content analysis 93%
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 94%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 94%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 94%
Similar papers in this journal
Similar papers in this journal
- Perceptions of Complementary, Alternative, and Integrative Medicine: Insights from a Large-Scale International Cross-Sectional Survey of Surgery Researchers and Clinicians 92%
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 91%
- Self-learning on COVID-19 among medical students and their preparedness to participate in government's COVID-19 response in Bhutan: a cross-sectional study 91%
Similar papers in this journal
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 94%
- Data Driven Monitoring in Community Based Management of SAM children using Psychometric Techniques: An Operational Framework 91%
- Validation of the patient reported outcome measures tool “Catquest” in Odia language 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.