AI in Medical Education: A Comparative Analysis of GPT-4 and GPT-3.5 on Turkish Medical Specialization Exam Performance
Kilic, M. E.
Show abstract
Background/aimLarge-scale language models (LLMs), such as GPT-4 and GPT-3.5, have demonstrated remarkable potential in the rapidly developing field of artificial intelligence (AI) in education. The use of these models in medical education, especially their effectiveness in situations such as the Turkish Medical Specialty Examination (TUS), is yet understudied. This study evaluates how well GPT-4 and GPT-3.5 respond to TUS questions, providing important insight into the real-world uses and difficulties of AI in medical education. Materials and methodsIn the study, 1440 medical questions were examined using data from six Turkish Medical Specialties examinations. GPT-4 and GPT-3.5 AI models were utilized to provide answers, and IBM SPSS 26.0 software was used for data analysis. For advanced enquiries, correlation analysis and regression analysis were used. ResultsGPT-4 demonstrated a better overall success rate (70.56%) than GPT-3.5 (40.17%) and physicians (38.14%) in this study examining the competency of GPT-4 and GPT-3.5 in answering questions from the Turkish Medical Specialization Exam (TUS). Notably, GPT-4 delivered more accurate answers and made fewer errors than GPT-3.5, yet the two models skipped about the same number of questions. Compared to physicians, GPT-4 produced more accurate answers and a better overall score. In terms of the number of accurate responses, GPT-3.5 performed slightly better than physicians. Between GPT-4 and GPT-3.5, GPT-4 and the doctors, and GPT-3.5 and the doctors, the success rates varied dramatically. Performance ratios differed across domains, with doctors outperforming AI in tests involving anatomy, whereas AI models performed best in tests involving pharmacology. ConclusionsIn this study, GPT-4 and GPT-3.5 AI models showed superior performance in answering Turkish Medical Specialization Exam questions. Despite their abilities, these models demonstrated limitations in reasoning beyond given knowledge, particularly in anatomy. The study recommends adding AI support to medical education to enhance the critical interaction with these technologies.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- Large language models for generating medical examinations: systematic review 97%
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 96%
- Medical students' perceptions towards artificial intelligence in education and practice: A multinational, multicenter cross-sectional study 96%
Similar papers in this journal
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 96%
- High School Science Fair: What Students Say -- Mastery, Performance, and Self-Determination Theory 95%
- A national professional development program fills mentoring gaps for postdoctoral researchers 95%
Similar papers in this journal
- Evaluation of Self-Directed Learning Activities at King Abdulaziz University: A Qualitative Study of Faculty Perceptions 95%
- Validation of the patient reported outcome measures tool “Catquest” in Odia language 90%
- Data Driven Monitoring in Community Based Management of SAM children using Psychometric Techniques: An Operational Framework 89%
Similar papers in this journal
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 93%
- Perceptions of Complementary, Alternative, and Integrative Medicine: Insights from a Large-Scale International Cross-Sectional Survey of Surgery Researchers and Clinicians 93%
- Self-learning on COVID-19 among medical students and their preparedness to participate in government's COVID-19 response in Bhutan: a cross-sectional study 92%
Similar papers in this journal
- Performance of o1 pro and GPT-4 in self-assessment questions for nephrology board renewal 94%
- For-Profit Growth and Academic Decline: A Retrospective Nationwide Assessment of Brazilian Medical Schools 92%
- Emerging Applications of NLP and Large Language Models in Gastroenterology and Hepatology: A Systematic Review 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.