Don't stop the heart: a performance analysis of large language models and potassium dosing
Blotske, K.; Zhao, X.; Henry, K.; Murray, B.; Gao, Y.; Smith, S. E.; Wayne, N.; Ku, P.; Smith, B.; Moua, S.; Sikora, A.
Show abstract
Background: Electrolyte replacement is ubiquitous in the acute care setting, but its familiarity cannot belie that even small dosing errors with potassium can cause lethal cardiac arrhythmias. Recently, MedAgentBench offered a benchmark for agentic artificial intelligence (AI) including the ability to correctly dose potassium based on a single rule; however, this does not adequately reflect the clinical complexity or safety concerns of an agent that has been used as the lethal injection. The purpose of this analysis was to a probe leaderboard large language model (LLM) capabilities to follow basic dosing rules to safely replace potassium in a series of clinician-annotated cases. Methods: Using a clinician panel, we developed a series of dosing principles and 20 clinical cases reflective of the complexity of potassium replacement. External clinicians were surveyed to assess practice variability and agreement to clinician panel answers. We tested GPT-5-chat with each case in triplicate, with and without the clinician curated dosing principles, and prompted the model to answer six questions involving potassium goals, dosing, route, lab frequency, concurrent interventions, and the model's perceived level of confidence for the output and complexity of the case. The primary outcome was the rate of appropriate recommendations in comparison to clinician answers. Results: A total of 54 clinicians reviewed the 20 hypokalemia cases and hypokalemia dosing guideline. Clinicians expressed "highly agree" or "somewhat agree" for 66.8% of the cases evaluated when asked if they agree with the guideline-recommended management. When given the potassium dosing guideline, total errors dropped from 165 to 104, and average accuracy improved from 45% to 65% with GPT-5-Chat. GPT-5-Chat conveyed a high level of confidence for 100% of responses, while labeling 80% and 76% of cases as highly complex with and without the criteria, respectively. Potential harm scores were considerable in both groups, however, a notable reduction in severity scores occurred with the dosing guidance document. Recommendations on concurrent interventions and dosing had the highest rate of errors in both groups. Conclusions: Benchmarks must appropriately reflect clinical complexity to be considered valuable for the deployment of agentic artificial intelligence tools in the healthcare domain. GPT-5-Chat assessment on a comprehensive medication management task for potassium replacement showed improvement with dosing guidance, yet unfit benchmarking performance.
Matching journals
The top 10 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Medication Clusters at Hospital Discharge and Risk of Adverse Drug Events at 30-days Post-Discharge: A Population-based Cohort Study of Older Adults 91%
- Sensitivity of Estimated Tacrolimus Population Pharmacokinetic Profile to Inaccurate Assumptions about Dose Timing and Absorption: An Investigation in Real-World and Simulated Data 90%
- Population Pharmacokinetic Analysis of Dexmedetomidine in Children using Real-World-Data Obtained from Electronic Health Records and Remnant Clinic Samples 89%
Similar papers in this journal
- MTXPK.org: A clinical decision support tool evaluating high-dose methotrexate pharmacokinetics to inform post-infusion care and use of glucarpidase 91%
- Algorithmic identification of treatment-emergent adverse events from clinical notes using large language models: a pilot study in inflammatory bowel disease 89%
- DrugWAS: Leveraging drug-wide association studies to facilitate drug repurposing for COVID-19 88%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Pharmacogenomics implementation training improve self-efficacy and competency to drive adoption in clinical practice 94%
- PBPK modelling of dexamethasone in patients with COVID-19 and liver disease 89%
- A systematic review of utilisation of diurnal timing information in clinical trial design for Long QT syndrome 88%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.