Improved Performance of ChatGPT-4 on the OKAP Exam: A Comparative Study with ChatGPT-3.5
Teebagy, S.; Colwell, L.; Wood, E.; Yaghy, A.; Faustina, M.
Show abstract
This study aims to evaluate the performance of ChatGPT-4, an advanced Artificial Intelligence (AI) language model, on the Ophthalmology Knowledge Assessment Program (OKAP) examination compared to its predecessor, ChatGPT-3.5. Both models were tested on 180 OKAP practice questions covering various ophthalmology subject categories. Results showed that ChatGPT-4 significantly outperformed ChatGPT-3.5 (81% vs. 57%; p<0.001), indicating improvements in medical knowledge assessment. The superior performance of ChatGPT-4 suggests potential applicability in ophthalmologic education and clinical decision support systems. Future research should focus on refining AI models, ensuring a balanced representation of fundamental and specialized knowledge, and determining the optimal method of integrating AI into medical education and practice.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- VisualR: a novel and scalable solution for assessing visual function using virtual reality 91%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 91%
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 91%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- Towards implementation of AI in New Zealand national screening program: Cloud-based, Robust, and Bespoke 92%
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.