Evaluating Large Language Models for Drafting Emergency Department Discharge Summaries
Williams, C. Y. K.; Bains, J.; Tang, T.; Patel, K.; Lucas, A. N.; Chen, F.; Miao, B. Y.; Butte, A. J.; Kornblith, A. E.
Show abstract
ImportanceLarge language models (LLMs) possess a range of capabilities which may be applied to the clinical domain, including text summarization. As ambient artificial intelligence scribes and other LLM-based tools begin to be deployed within healthcare settings, rigorous evaluations of the accuracy of these technologies are urgently needed. ObjectiveTo investigate the performance of GPT-4 and GPT-3.5-turbo in generating Emergency Department (ED) discharge summaries and evaluate the prevalence and type of errors across each section of the discharge summary. DesignCross-sectional study. SettingUniversity of California, San Francisco ED. ParticipantsWe identified all adult ED visits from 2012 to 2023 with an ED clinician note. We randomly selected a sample of 100 ED visits for GPT-summarization. ExposureWe investigate the potential of two state-of-the-art LLMs, GPT-4 and GPT-3.5-turbo, to summarize the full ED clinician note into a discharge summary. Main Outcomes and MeasuresGPT-3.5-turbo and GPT-4-generated discharge summaries were evaluated by two independent Emergency Medicine physician reviewers across three evaluation criteria: 1) Inaccuracy of GPT-summarized information; 2) Hallucination of information; 3) Omission of relevant clinical information. On identifying each error, reviewers were additionally asked to provide a brief explanation for their reasoning, which was manually classified into subgroups of errors. ResultsFrom 202,059 eligible ED visits, we randomly sampled 100 for GPT-generated summarization and then expert-driven evaluation. In total, 33% of summaries generated by GPT-4 and 10% of those generated by GPT-3.5-turbo were entirely error-free across all evaluated domains. Summaries generated by GPT-4 were mostly accurate, with inaccuracies found in only 10% of cases, however, 42% of the summaries exhibited hallucinations and 47% omitted clinically relevant information. Inaccuracies and hallucinations were most commonly found in the Plan sections of GPT-generated summaries, while clinical omissions were concentrated in text describing patients Physical Examination findings or History of Presenting Complaint. Conclusions and RelevanceIn this cross-sectional study of 100 ED encounters, we found that LLMs could generate accurate discharge summaries, but were liable to hallucination and omission of clinically relevant information. A comprehensive understanding of the location and type of errors found in GPT-generated clinical text is important to facilitate clinician review of such content and prevent patient harm.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
Similar papers in this journal
- Enhancing Research Data Infrastructure to Address the Opioid Epidemic: The Opioid Overdose Network (02-Net) 94%
- A Machine Learning Approach to Identifying Delirium from Electronic Health Records 91%
- Adopting an American framework to optimize nursing admission documentation in an Australian health organization 91%
Similar papers in this journal
- Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions: A National Retrospective EHR Study 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 92%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 92%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 91%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.