ChatGPT as a Digital Pharmacist: A Systematic Review and Meta-Analysis of Drug-Counselling Accuracy
Azmakan, H.; Nabipour, A.; Ghorabi Tehrani, N.; Najari, N.; Fathi hafshjani, P.; Falahati Marvast, A.; Mani, S.; Asemi Sichani, N.; Fallah Pakdaman, S.; Shieh, M.; Afrandkhalilabad, Z.; Shadi, A.; Shahidi, R.
Show abstract
BackgroundThe emergence of Large Language Models (LLMs) like ChatGPT presents significant opportunities for healthcare, yet raises concerns about accuracy, especially in high-risk areas such as medication counseling. A comprehensive evaluation of ChatGPTs reliability in providing drug information is crucial for its safe integration into clinical practice. This systematic review and meta-analysis aimed to assess the accuracy of drug-counseling information provided by ChatGPT 4. MethodsFollowing PRISMA, we systematically searched PubMed, Embase, Scopus, and Web of Science on May 9, 2025, for original research evaluating the accuracy of ChatGPT (version 4 or newer) in drug-counseling queries. Included studies compared the AIs output against standard comparators like pharmacists or drug databases. A random-effects meta-analysis was performed to calculate the pooled proportion of accurate responses, and study quality was assessed using a customized Newcastle-Ottawa Scale (NOS). ResultsThe search identified 17 eligible studies. Of these, 15 were included in the meta-analysis, which showed a pooled accuracy rate of 86% (95% CI: 0.75-0.95). However, significant heterogeneity was observed across studies (I2=98.5%, p<0.0001). Quality of the studies was a concern, with only four studies (24%) rated as high quality. No evidence of publication bias was found (p=0.91). ConclusionChatGPT demonstrates substantial promise in drug counseling, with an 86% accuracy rate that surpasses its performance in other medical domains. However, the high heterogeneity and a non-trivial 14% error rate, coupled with methodological weaknesses in the primary literature, indicate that ChatGPT is not yet ready for autonomous clinical use. Its current role should be as a supplementary tool under the strict supervision of qualified healthcare professionals to ensure patient safety.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- To Include or Not to Include? A prescription from the pharmacy on how to use active learning assisted screening in systematic reviews 95%
- Is artificial intelligence for medical professionals serving the patients? Protocol for a mixed method systematic review on patient-relevant benefits and harms of algorithmic decision-making 93%
- Repurposing Existing Medications for Coronavirus Disease 2019: Protocol for a Rapid and Living Systematic Review 93%
Similar papers in this journal
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 93%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 92%
Similar papers in this journal
- Application of observational research methods to real-world studies for rare disease drugs: a scoping review protocol 95%
- Modelling the impact of behavioural interventions during pandemics: A systematic review 94%
- Common misconceptions held by health researchers when interpreting linear regression assumptions, a cross-sectional study 94%
Similar papers in this journal
- Pharmacogenomics implementation training improve self-efficacy and competency to drive adoption in clinical practice 94%
- Real-world observation on response to cholinesterase inhibitors or selective serotonin reuptake inhibitors prescribed to outpatients with dementia using electronic medical records 91%
- Assessment of Hydroxychloroquine and Chloroquine Safety Profiles: A Systematic Review and Meta-Analysis 90%
Similar papers in this journal
- Patient-Reported Reasons for Antihypertensive Medication Change: A Quantitative Study Using Social Media 95%
- Standardization of drug names in the FDA Adverse Event Reporting System: The DiAna dictionary 95%
- Large-scale empirical identification of candidate comparators for pharmacoepidemiological studies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.