Augmentation of ChatGPT with Clinician-Informed Tools Improves Performance on Medical Calculation Tasks
Goodell, A. J.; Chu, S. N.; Rouholiman, D.; Chu, L. F.
Show abstract
AO_SCPLOWBSTRACTC_SCPLOWPrior work has shown that large language models (LLMs) have the ability to answer expert-level multiple choice questions in medicine, but are limited by both their tendency to hallucinate knowledge and their inherent inadequacy in performing basic mathematical operations. Unsurprisingly, early evidence suggests that LLMs perform poorly when asked to execute common clinical calculations. Recently, it has been demonstrated that LLMs have the capability of interacting with external programs and tools, presenting a possible remedy for this limitation. In this study, we explore the ability of ChatGPT (GPT-4, November 2023) to perform medical calculations, evaluating its performance across 48 diverse clinical calculation tasks. Our findings indicate that ChatGPT is an unreliable clinical calculator, delivering inaccurate responses in one-third of trials (n=212). To address this, we developed an open-source clinical calculation API (openmedcalc.org), which we then integrated with ChatGPT. We subsequently evaluated the performance of this augmented model by comparing it against standard ChatGPT using 75 clinical vignettes in three common clinical calculation tasks: Caprini VTE Risk, Wells DVT Criteria, and MELD-Na. The augmented model demonstrated a marked improvement in accuracy over unimproved ChatGPT. Our findings suggest that integration of machine-usable, clinician-informed tools can help alleviate the reliability limitations observed in medical LLMs.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 96%
- Development and Validation of Phenotype Classifiers across Multiple Sites in the Observational Health Sciences and Informatics (OHDSI) Network 94%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 94%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 95%
- Evaluating Semantic Similarity Methods for Comparison of Text-derived Phenotype Profiles 93%
- Optimized Feature Selection and Advanced Machine Learning for Stroke Risk Prediction in Revascularized Coronary Artery Disease Patients 92%
Similar papers in this journal
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 94%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 94%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.