Rx-LLM: a benchmarking suite to evaluate safe LLM performance for medication-related tasks
Zhao, X.; Blotske, K.; Cargile, M.; Tilley, A.; Murray, B.; Gao, Y.; Henry, K.; Smith, S.; Barreto, E.; Bauer, S.; Sohn, S.; Liu, T.; Sikora, A.
Show abstract
BackgroundFor large language models (LLMs) to reach their potential as information technology tools that make medication use safer, clinically relevant benchmarks capable of automated grading and designed specifically to measure the performance of LLMs for medication tasks are required. The purpose of this study was to design a suite of benchmarking tests reflective of Comprehensive Medication Management (CMM; the standard of care for medication optimization) and quantify the baseline performance of the latest LLMs. MethodsWe established six benchmarks representing critical stages of the CMM process: drug formulation matching, drug order (sig) generation, drug route matching, drug-drug interaction identification, renal dose identification, and drug-indication matching. For each benchmark, we curated a clinician-annotated dataset comprising 250 standardized input-output pairs including both inpatient and outpatient medications. We evaluated the clinical knowledge retrieval capabilities of three LLMs: GPT-4o-mini, MedGemma-27B, and LLaMA3-70B. We employed a zero-shot prompting strategy, excluding in-context examples, to assess the models internal clinical knowledge rather than their few-shot learning potential. To check reliability, each model was run three times using a temperature of 0.7 (a mid-range value of an LLM setting controlling text generation randomness). Performance was assessed using task-specific evaluation metrics including precision (positive predictive value), recall (sensitivity), F1-score, accuracy, and correctness consistency across trials. ResultsAcross six benchmarks, LLaMA3-70B demonstrated the highest performance in four tasks: drug-formulation matching (F1, 54.0% [95 CI: 50.1-58]), drug-order generation (accuracy, 88.0%), drug-route identification (F1, 74.3% [95 CI: 71-78]), and drug-indication identification (accuracy, 97.6% [95 CI: 95.6-99.2]). In the drug-drug interaction task, GPT-4o-mini achieved the highest overall accuracy (70.4% [95 CI: 64.8-75.7]). For renal dose-adjustment identification, GPT-4o-mini demonstrated the highest F1 score (83.3% [95 CI: 77.6-88]). Correctness-consistency scores ranged from 8.0% to 97.6% across benchmarks, with no model exhibiting uniformly superior consistency. ConclusionsModel performance varied substantially across medication-related tasks. LLaMA3-70B demonstrated promising baseline performance in tasks involving formulation, ordering, route, and indication. GPT-4o-mini showed potential advantages in drug-drug interaction detection and renal dose adjustment. These findings underscore the need for task-specific evaluation when deploying models for medication-focused clinical decision support.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 94%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 93%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 93%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 95%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 92%
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 91%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 91%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 91%
Similar papers in this journal
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 95%
- Biomedical Text Normalization through Generative Modeling 94%
- PK-RNN-V E: A Deep Learning Model Approach to Vancomycin Therapeutic Drug Monitoring Using Electronic Health Record Data 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.