A medically-grounded LLM agent-based tool to detect patient safety events in medical records
Trujillo, D.; Wang, D.; Bahr, N.; Yi-Jin Hsieh, T.; Cho, B.; Meckler, G.; Hansen, M.; Eriksson, C.; Seo Kim, K.; Bedrick, S.; jiang, x.; Guise, J.-M.
Show abstract
Large language models (LLMs) have shown incredible promise in medicine. While LLMs may be particularly useful in areas requiring extensive review of clinical records, their use remains limited due to their tendency to hallucinate and fabricate information. Hallucination issues, as well as their consequences, are exacerbated in low-probability, high-stakes scenarios such as rare adverse safety events or medical errors. We present SAFE-AI (Structured and Automated Framework for Explainable AI), a novel method for clinical decision making that combines the strengths of clinical expert knowledge with LLMs in an ontology-driven model that minimizes hallucinations using strict rules. We test this method to identify medication errors in medical charts. We collected a sample of 18,402 lines of clinical information from 300 EMS clinical charts that were independently dually reviewed by two expert physicians for epinephrine adverse safety events (ASEs), with 96% inter-rater agreement. We tested SAFE-AI against these labels, achieving human-like performance in detecting epinephrine overdoses with 97.9% accuracy, and 91.6% accuracy in identifying delays in epinephrine administration, greatly outperforming baseline LLMs models. Notably, some disagreements between clinicians and the model were found to be justifiable differences in judgment rather than errors. SAFE-AI presents a novel approach for clinical AI applications that addresses two key limitations of current machine learning methods: 1) over-reliance on probabilistic pattern recognition instead of established medical knowledge, and 2) perpetuation of biases present in training data. This framework is easily adaptable to a range of clinical applications, paving the way for provable and trustworthy AI in medicine. Author SummaryLLMs have shown promise in analyzing clinical records but their use is limited due to their tendency to hallucinate and fabricate information. Misinformation could threaten patient safety and jeopardize trust. We developed SAFE-AI (Structured and Automated Framework for Explainable AI), which combines knowledge from clinical experts with LLM inference to detect adverse safety events (ASEs) with minimal errors. We tested our method in identifying medication errors in medical charts and compared results to reviews by expert physicians. Our method detected epinephrine delays and overdoses with a high level of accuracy. SAFE-AI presents a novel approach for clinical AI applications that overcomes reliance on pattern recognition instead of medical knowledge biases present in training data.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automated stratification of trauma injury severity across multiple body regions using multi-modal, multi-class machine learning models 95%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 95%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 95%
Similar papers in this journal
- Modeling physician variability to prioritize relevant medical record information 95%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 95%
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 94%
Similar papers in this journal
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 95%
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 94%
- Causal Analysis for Multivariate Integrated Clinical and Environmental Exposures Data 93%
Similar papers in this journal
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 95%
- Signal from the Noise: A Mixed Methods Process Mining Approach to Evaluate Care Pathways. 95%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 94%
Similar papers in this journal
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 94%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 93%
- From theoretical models to practical deployment: A perspective and case study of opportunities and challenges in AI-driven healthcare research for low-income settings 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.