Testing new versions of ChatGPT in terms of physiology and electrophysiology of hearing: improved accuracy but not consistency
Jedrzejczak, W. W.; Skarzynski, H.; Kochanek, K.
Show abstract
IntroductionChatGPT has revolutionized many aspects of modern life, including scientific ones. Since its introduction, new versions have been introduced and advertised as having better performance. But is this true? This study aimed to assess the accuracy and consistency of six versions of ChatGPT (3.5, 4, 4o mini, 4o, 4o1 mini, and 4o1 preview). Of interest was the variability of responses given to asking the same question multiple times. MethodsWe evaluated 6 versions of ChatGPT based on their responses to 30 single-answer, multiple-choice exam questions from a 1-year course on objective methods of testing hearing. The questions were posed 10 times to each version of ChatGPT across two days (5 times each day). The accuracy of the responses was evaluated in terms of a response key. To evaluate consistency (repeatability) of the responses over time, percent agreement and Cohens Kappa were calculated. ResultsThe overall accuracy of ChatGPT increased with each version, starting from around 53% for version 3.5 and rising to 86% for version 4o1 preview. The greatest improvement in accuracy and repeatability came with the introduction of version 4o. Repeatability progressively rose with newer releases with the exception of version 4o1 mini. While the current top version 4o1 preview has similar repeatability to 4o, the faster version, 4o1 mini, had significantly lower repeatability than the older 4o mini. ConclusionNewer versions of ChatGPT generally show improvement in terms of accuracy, but not in repeatability. The variability of responses is probably the current main limitation of ChatGPT for professional applications. Users must be especially careful with version 4o1 mini.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cross-modal sensory boosting to improve high-frequency hearing loss 93%
- Challenges in Implementing a Mobile AI Chatbot Intervention for Depression Among Youth on Psychiatric Waiting Lists: A Randomized Control Study Termination Report 90%
- Prediction of COVID-19 Mortality to Support Patient Prognosis and Triage and Limits of Current Open-Source Data 89%
Similar papers in this journal
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 94%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
Similar papers in this journal
- Performance of ChatGPT in pediatric audiology as rated by students and experts 97%
- Feasibility of an App-Assisted and Home-Based Video Version of the Timed up and Go Test for Patients with Parkinson Disease: vTUG 90%
- Artificial Intelligence in laryngeal endoscopy: Systematic Review and Meta-Analysis 89%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.