Assessing the Utility of ChatGPT Throughout the Entire Clinical Workflow
Rao, A. S.; Pang, M.; Kim, J.; Kamineni, M.; Lie, W.; Prasad, A. K.; Landman, A.; Dryer, K.; Succi, M. D.
Show abstract
IMPORTANCELarge language model (LLM) artificial intelligence (AI) chatbots direct the power of large training datasets towards successive, related tasks, as opposed to single-ask tasks, for which AI already achieves impressive performance. The capacity of LLMs to assist in the full scope of iterative clinical reasoning via successive prompting, in effect acting as virtual physicians, has not yet been evaluated. OBJECTIVETo evaluate ChatGPTs capacity for ongoing clinical decision support via its performance on standardized clinical vignettes. DESIGNWe inputted all 36 published clinical vignettes from the Merck Sharpe & Dohme (MSD) Clinical Manual into ChatGPT and compared accuracy on differential diagnoses, diagnostic testing, final diagnosis, and management based on patient age, gender, and case acuity. SETTINGChatGPT, a publicly available LLM PARTICIPANTSClinical vignettes featured hypothetical patients with a variety of age and gender identities, and a range of Emergency Severity Indices (ESIs) based on initial clinical presentation. EXPOSURESMSD Clinical Manual vignettes MAIN OUTCOMES AND MEASURESWe measured the proportion of correct responses to the questions posed within the clinical vignettes tested. RESULTSChatGPT achieved 71.7% (95% CI, 69.3% to 74.1%) accuracy overall across all 36 clinical vignettes. The LLM demonstrated the highest performance in making a final diagnosis with an accuracy of 76.9% (95% CI, 67.8% to 86.1%), and the lowest performance in generating an initial differential diagnosis with an accuracy of 60.3% (95% CI, 54.2% to 66.6%). Compared to answering questions about general medical knowledge, ChatGPT demonstrated inferior performance on differential diagnosis ({beta}=-15.8%, p<0.001) and clinical management ({beta}=-7.4%, p=0.02) type questions. CONCLUSIONS AND RELEVANCEChatGPT achieves impressive accuracy in clinical decision making, with particular strengths emerging as it has more clinical information at its disposal.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 94%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 92%
Similar papers in this journal
Similar papers in this journal
- Performance of Advanced Large Language Models (GPT-4o, GPT-4, Gemini 1.5 Pro, Claude 3 Opus) on Japanese Medical Licensing Examination: A Comparative Study 92%
- Synthetic Data Generation in Healthcare: A Scoping Review of reviews on domains, motivations, and future applications 92%
- A Deep Learning Method to Detect Opioid Prescription and Opioid Use Disorder from Electronic Health Records 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.