Back

Towards Inpatient Discharge Summary Automation via Large Language Models: A Multidimensional Evaluation with a HIPAA-Compliant Instance of GPT-4o and Clinical Expert Assessment

Osborne, T. G.; Abbasi, S. M.; Hong, S. M.; Sexton, R. T.; Ambut, J.; Patel, N. J.; Rosenthal, R. N.; Ung, L.; Wang, F.; Wong, R.

2025-04-04 health systems and quality improvement
10.1101/2025.04.03.25325204 medRxiv
Show abstract

Large language models (LLMs) have demonstrated potential to automate clinical documentation tasks that may reduce clinician burden, such as generation of hospital discharge summaries. Prior research used older LLMs and limited data, raising concerns about fabrications and omissions. In this study, we evaluated the automatic generation of inpatient Internal Medicine discharge summaries using a HIPAA-compliant Microsoft Azure instance of OpenAIs GPT-4o. Both human-written and AI-generated discharge summaries were scored by Internal Medicine hospital faculty for quality, readability/conciseness, factuality and completeness, presence of hallucinations/omissions and their impact on safety, and compared with the actual discharge summaries. Our results showed that the AI-generated discharge summaries significantly outperformed actual human written summaries in both quality and readability/conciseness and were comparable to humans in factuality and completeness, with a minimal cost.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.