Physician- versus Large Language Model-Generated Summaries in the Emergency Department
Golchini, N. B.; Mehandru, N.; Alaa, A.; Molina, M.
Show abstract
BackgroundAs part of routine practice and documentation, emergency department (ED) clinicians routinely construct "one-liner" summaries--brief, information-rich statements distilling a patients history and presentation to support rapid decision-making. Producing these summaries is cognitively demanding and contributes to documentation burden. Large language models (LLMs) may assist by synthesizing longitudinal electronic health record (EHR) data. MethodsWe conducted a blinded, within-subject study of 99 ED encounters from March 2022-March 2024 at the University of California, San Francisco. We used an LLM to generate one-liner summaries using a k-nearest-neighbor few-shot prompting approach and clinical notes spanning multiple prior encounters. Twenty-one emergency physicians evaluated paired LLM- and physician-authored summaries in randomized order, rating accuracy, completeness, and clinical utility on 5-point Likert scales and indicating their overall preference with optional free-text explanation. Ratings were analyzed using linear mixed-effects models with summary type as a fixed effect and reviewer as a random intercept. Secondary analyses examined the LLMs note-selection behavior and how inclusion of specific note types affected summary quality. We used rapid content analysis to review free-text explanations, identifying recurrent themes among reviewer preferences. ResultsAcross all dimensions, LLM-generated summaries received higher ratings than physician-authored summaries. Mean (SE) estimated marginal means for accuracy were 4.18 (0.09) vs 3.40 (0.11) ({beta} = 0.78; 95% CI 0.50-1.07; p < .001), for completeness 3.69 (0.10) vs 3.25 (0.12) ({beta} = 0.44; 95% CI 0.14-0.74; p = .005), and for clinical utility 3.88 (0.10) vs 3.21 (0.12) ({beta} = 0.67; 95% CI 0.35-0.99; p < .001). LLM-generated summaries were preferred in 50.5% of encounters, physician summaries in 38.4%, and 11.1% were rated equivalent. Qualitative analysis indicated that LLM summaries were often more inclusive and neutrally phrased, whereas physician summaries exhibited greater contextual nuance but occasionally omitted key details. ConclusionsIn this blinded evaluation of ED encounters, LLM-generated one-liner summaries outperformed physician-authored summaries on accuracy, completeness, and clinical utility. Patterns in note utilization suggest that models selectively integrate high-yield clinical sources, which may have important implications for the cost and efficiency in future healthcare deployment. These findings represent an important first step toward leveraging LLMs to aid rapid synthesis of complex EHR data in high-stakes environments.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
Similar papers in this journal
- Theory of radiologist interaction with instant messaging decision support tools: a sequential-explanatory study 94%
- Accuracy of preferred language data in a multi-hospital electronic health record in Toronto, Canada 93%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 93%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 93%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.