Towards Inpatient Discharge Summary Automation via Large Language Models: A Multidimensional Evaluation with a HIPAA-Compliant Instance of GPT-4o and Clinical Expert Assessment
Osborne, T. G.; Abbasi, S. M.; Hong, S. M.; Sexton, R. T.; Ambut, J.; Patel, N. J.; Rosenthal, R. N.; Ung, L.; Wang, F.; Wong, R.
Show abstract
Large language models (LLMs) have demonstrated potential to automate clinical documentation tasks that may reduce clinician burden, such as generation of hospital discharge summaries. Prior research used older LLMs and limited data, raising concerns about fabrications and omissions. In this study, we evaluated the automatic generation of inpatient Internal Medicine discharge summaries using a HIPAA-compliant Microsoft Azure instance of OpenAIs GPT-4o. Both human-written and AI-generated discharge summaries were scored by Internal Medicine hospital faculty for quality, readability/conciseness, factuality and completeness, presence of hallucinations/omissions and their impact on safety, and compared with the actual discharge summaries. Our results showed that the AI-generated discharge summaries significantly outperformed actual human written summaries in both quality and readability/conciseness and were comparable to humans in factuality and completeness, with a minimal cost.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development of a Post-Acute Sequelae of COVID-19 (PASC) Symptom Lexicon Using Electronic Health Record Clinical Notes 95%
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 95%
- EHR-QC: A streamlined pipeline for automated electronic health records standardisation and preprocessing to predict clinical outcomes 95%
Similar papers in this journal
- Evaluating the impact on clinical task efficiency of a natural language processing algorithm for searching medical documents: Prospective crossover study 97%
- Transformative potential of Large Language Models in data mining on Electronic Health Records. 95%
- Is the quality of hospital EHR data sufficient to evidence its ICHOM outcomes performance in heart failure? A pilot evaluation 93%
Similar papers in this journal
- Evaluating Anti-LGBTQIA+ Medical Bias in Large Language Models 94%
- From months to minutes: creating Hyperion, a novel data management system expediting data insights for oncology research and patient care 94%
- Ethical review of clinical research with generative AI: Evaluating ChatGPT’s accuracy and reproducibility 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.