MedMatch: a first step for the automation of large language model performance benchmarking for medication-related tasks
Blotske, K.; Zhao, X.; Cargile, M.; Tilley, A.; Murray, B.; Gao, Y.; Henry, K.; Smith, S. E.; Barreto, E.; Bauer, S.; Sohn, S.; Liu, T.; Sikora, A.
Show abstract
BackgroundThe accuracy and safety of generating medication orders by large language models (LLMs) must be demonstrated. Without standardization, performance evaluation is limited to time and resource-intensive clinician grading. This evaluation aimed to develop a standardized medication format that supports automated performance evaluation (MedMatch). MethodsFirst, a survey of 40 medication prompts was given to clinicians to assess agreement in medication order communication. Second, a clinician panel developed a standardized medication format (MedMatch) for oral and intravenous medications. Third, a clinician-annotated dataset of medication prompts and standardized answers in the MedMatch format was developed for LLM testing. Finally, LLMs were retested with the same dataset, adjusted to exclude route information, to evaluate the appropriate categorization of medication route. ResultsThe formal medication orders consistently showed low omission rates and high overlap for all entities, compared to the verbal and brief written communication types. Lexical overlap results demonstrated pattern norms amongst clinicians with entities appearing most commonly in positions 1-5 in the order of drug name, dose, unit, route, and frequency. In the second survey, the formal written group performed the highest with 78.3% of prompts considered appropriate as a computer-generated response. LLM accuracy on MedMatch order standardization was highest in oral solid (64.2-72.5%), intravenous intermittent (72.5-84.3%), and intravenous push (62.7-74.5%) categories. LLMs performed the worst at categorizing medication orders accurately into intravenous push (18-61%) and intravenous intermittent (51-100%) routes. ConclusionsStandardized format for computer-based outputs may support automated performance analysis and enhance the clarity of medication communication.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Determining prescriptions in electronic health care (EHR) data: methods for development of standardised, reproducible drug codelists 94%
- Enhancing Research Data Infrastructure to Address the Opioid Epidemic: The Opioid Overdose Network (02-Net) 93%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 93%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 94%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 93%
Similar papers in this journal
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 91%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 91%
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 91%
Similar papers in this journal
- Impact of CancelRx on Discontinuation of Controlled Substance Prescriptions 95%
- Electronic prescribing systems as tools to improve patient care: a learning health systems approach to increase guideline concordant prescribing for venous thromboembolism prevention 91%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.