Proposed Context-of-Use Evaluation Framework for Medication Management Tasks Completed by Generative Artificial Intelligence
Henry, K.; Blotske, K.; Smith, B.; Li, T.; Gao, Y.; Zhao, X.; Liu, T.; Sikora, A.
Show abstract
Background: Standardized evaluation of agentic artificial intelligence (AI) for medication management is lacking. Given the potential lethality of medication errors endorsed or missed by AI, performance evaluation constructs are essential. The purpose of this evaluation was to develop a standardized grading framework for performance evaluation of medication management tasks. Methods: A mixed-methods approach was undertaken that included literature evaluation for standards and best practices of comprehensive medication management (CMM), panel discussions, and iterative application to set of cases. The goal was to develop a grading framework that effectively evaluated domains like safety, factuality, and clinical relevance that can be employed for a broad range of medication domains (i.e., electrolyte replacement, antibiotic selection). Inter-rater reliability with intraclass Krippendorffs Alpha was the primary outcome. Results: A total of 5 panelists developed the CMM Evaluation Framework, which includes 4 dimensions: safety, factuality, completeness, and preference. These dimensions are applied to three CMM skills: collecting patient data, analyzing information, and designing regimens. Each dimension is rated from 1-5. An additional dimension evaluated the presence of hallucinations and errors with high harm scores (i.e., absolute failure criteria regardless of an overall score). The Krippendorffs Alpha was highest in the medication therapy problem and medication therapy format categories, for 50 pneumonia cases, run in triplicate (150 total). Conclusions: This framework is informed by national standards for CMM and the healthcare professionals dedicated to the provision of this service. These domains allow for the possibilities of practice variation via the preference domain while also having strong guardrails against the commission of medication errors. Further analyses beyond pilot testing are necessary.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Active vaccine safety surveillance via a scalable, integrated system in Australian pharmacies 91%
- A Rapid Realist Review of the Role of Community Pharmacy in the Public Health Response to COVID-19 90%
- Perceptions, views and practices regarding antibiotic prescribing and stewardship among hospital physicians in Jakarta, Indonesia 90%
Similar papers in this journal
Similar papers in this journal
- A qualitative study on factors influencing health workers’ uptake of a pilot surgical antibiotic prophylaxis stewardship programme in selected Georgian hospitals 91%
- Healthcare Consumers’ Perceptions of Incentive-Linked Prescribing: A Scoping Review of Research 91%
- Acceptability of fixed-dose combination treatments for hypertension in Kenya: a qualitative study using the Theoretical Framework of Acceptability 90%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 93%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 92%
- Clinical Implementation Of Preemptive Pharmacogenomics Testing For Personalized Medicine At An Academic Medical Center 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.