Clinical Knowledge and Reasoning Abilities of AI Large Language Models in Pharmacy: A Comparative Study on the NAPLEX Exam
Angel, M.; Patel, A.; Alachkar, A.; Baldi, P. F.
Show abstract
ObjectiveThis study aims to evaluate the capabilities and limitations of three large language models (LLMs) - GPT-3, GPT-4, and Bard, in the field of pharmaceutical sciences by assessing their pharmaceutical reasoning abilities on a sample North American Pharmacist Licensure Examination (NAPLEX). We also analyze the potential impacts of LLMs on pharmaceutical education and practice. MethodsA sample NAPLEX exam consisting of 137 multiple-choice questions was obtained from an online source. GPT-3, GPT-4, and Bard were used to answer the questions by inputting them into the LLMs user interface. The answers provided by the LLMs were then compared with the answer key. ResultsGPT-4 exhibited superior performance compared to GPT-3 and Bard, answering 78.8% of the questions correctly. This score was 11% higher than Bard and 27.7% higher than GPT-3. However, when considering questions that required multiple selections, the performance of each LLM decreased significantly. GPT-4, GPT-3, and Bard only correctly answered 53.6%, 13.9%, and 21.4% of these questions, respectively. ConclusionAmong the three LLMs evaluated, GPT-4 was the only model capable of passing the NAPLEX exam. Nevertheless, given the continuous evolution of LLMs, it is reasonable to anticipate that future models will effortlessly pass the exam. This highlights the significant potential of LLMs to impact the pharmaceutical field. Hence, we must evaluate both the positive and negative implications associated with the integration of LLMs in pharmaceutical education and practice.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The Mastery Rubric for Bioinformatics: supporting design and evaluation of career-spanning education and training 92%
- Automated recognition of functional compound-protein relationships in literature 91%
- Assessment of research ethics education offerings of pharmacy master programs: a qualitative content analysis 91%
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 94%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 93%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
Similar papers in this journal
Similar papers in this journal
- Quantifying Device Type and Handedness Biases in a Remote Parkinson’s Disease AI-Powered Assessment 90%
- Utilization of Generative AI-drafted Responses for Managing Patient-Provider Communication 90%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 89%
Similar papers in this journal
- Cloud-controlled microscopy enables remote project-based biology education in Latinx communities in the United States and Latin America 90%
- Perceptions of Complementary, Alternative, and Integrative Medicine: Insights from a Large-Scale International Cross-Sectional Survey of Surgery Researchers and Clinicians 89%
- Linear vector models of time perception account for saccade and stimulus novelty interactions 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.