Generative Large Language Models in Electronic Health Records for Patient Care Since 2023: A Systematic Review
Du, X.; Wang, Y.; Zhou, Z.; Chuang, Y.-W.; Yang, R.; Zhang, W.; Wang, X.; Zhang, R.; Hong, P.; Bates, D. W.; Zhou, L.
Show abstract
BackgroundThe use of generative large language models (LLMs) with electronic health record (EHR) data is rapidly expanding to support clinical and research tasks. This systematic review synthesizes current strategies, challenges, and future directions for adapting and evaluating generative LLMs in EHR analyses and applications. MethodsWe followed the PRISMA guidelines to conduct a systematic review of articles from PubMed and Web of Science published between January 1, 2023, and November 9, 2024. Studies were included if they used generative LLMs to analyze real-world EHR data and reported quantitative performance evaluations. Through data extraction, we identified clinical specialties and tasks for each included article, and summarized evaluation methods. ResultsOf the 18,735 articles retrieved, 196 met our criteria. Most studies focused on Radiology (26.0%), Oncology (10.7%), and Emergency Medicine (6.6%). Regarding clinical tasks clinical decision support has the most studies of 62.2%, while summarizations and patient communications have the least studies of 5.6% and 5.1% separately. In addition, GPT-4 and ChatGPT were mostly used generative LLMs, which were used in 60.2% and 57.7% of studies, respectively. We identified 22 unique non-NLP metrics and 35 unique NLP metrics. Although NLP metrics have better scalability, none of the metrics were identified as having a strong correlation with gold-standard human evaluations. ConclusionOur findings highlight the need to evaluate generative LLMs on EHR data across a broader range of clinical specialties and tasks, as well as the urgent need for standardized, scalable, and clinically meaningful evaluation frameworks.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 96%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.