Evaluating Accuracy and Reproducibility of Large Language Model Performance in Pharmacy Education
Most, A.; MRC-ICU Investigator Team, ; Sikora, A.
Show abstract
The purpose of this study was to compare performance of ChatGPT (GPT-3.5), ChatGPT (GPT-4), Claude2, Llama2-7b, and Llama2-13b on 219 multiple-choice questions focusing on critical care pharmacotherapy. To further assess the ability of engineering LLMs to improve reasoning abilities and performance, we examined responses with a zero-shot Chain-of-Thought (CoT) approach, CoT prompting, and a custom built GPT (PharmacyGPT). A 219 multiple-choice questions focused on critical care pharmacotherapy topics used in Doctor of Pharmacy curricula from two accredited colleges of pharmacy was compiled for this study. A total of five LLMs were evaluated: ChatGPT (GPT-3.5), ChatGPT (GPT-4), Claude2, Llama2-7b, and Llama2-13b. The primary outcome was response accuracy. Of the five LLMs tested, GPT-4 showed the highest average accuracy rate at 71.6%. A larger variance indicates lower consistency and reduced confidence in its answers. Llama2-13b had the lowest variance (0.070) of all the LLMs, but performed with an accuracy of 41.5%. Following analaysis of overall accuracy, performance on knowledge- vs. skill-based questions were assessed. All five LLMs demonstrated higher accuracy on knowledge-based questions compared to skill-based questions. GPT-4 had the highest accuracy for knowledge- and skill-based questions, with an accuracy of 87% and 67%, respectively. Response accuracy from LLMs in the domain of clinical pharmacy can be improved by using prompt engineering techniques.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of ChatGPT on Chinese National Medical Licensing Examinations: A Five-Year Examination Evaluation Study for Physicians, Pharmacists and Nurses 95%
- Large language models for generating medical examinations: systematic review 93%
- Training Doctoral Students in Critical Thinking and Experimental Design using Problem-based Learning 92%
Similar papers in this journal
- Performance of Advanced Large Language Models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese Medical Licensing Examination: A Comparative Study 92%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 91%
- Development and Evaluation of MADDIE: Method to Acquire Delivery Date Information from Electronic Health Records 89%
Similar papers in this journal
- Assessment of research ethics education offerings of pharmacy master programs: a qualitative content analysis 92%
- Introducing the 4Ps Model of Transitioning to Distance Learning: a convergent mixed methods study conducted during the COVID-19 pandemic 92%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
Similar papers in this journal
- Development and Validation of a Machine Learning Model Integrated with the Clinical Workflow for Inpatient Discharge Date Prediction 90%
- Medical Clinical Minds Meet Artificial Intelligence: Italian Physicians' Knowledge, Attitudes, and Concordance between Italian Physicians and AI-Generated Diagnoses. A National Cross-Sectional Study 89%
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 88%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 92%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 90%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.