Achieving Expert-Level Clinical Infection Detection with LLMs from Clinical Documents: Validation in Complex Patient Cases with Cirrhosis
Yu, Y.; Ford, J.; Kim, E.; Patel, A.; Wardi, G.; Loomba, R.; Malhotra, A.; Nemati, S.; Ahn, J. C.
Show abstract
BackgroundSystemic infections are a leading cause of hospitalization and death among patients with cirrhosis. Timely and accurate infection identification is essential for both clinical care and the development of predictive models. However, existing methods such as ICD-10 coding are unreliable, and manual chart review is resource-intensive and difficult to scale. This study aimed to develop and validate an automated large language model (LLM)-based approach for infection classification and subtyping in patients with cirrhosis presenting to the emergency department (ED). MethodWe developed INFEHR (INfection identification and subtyping using Free-text EHR analysis), an LLM-powered pipeline utilizing Claude 3.5 Sonnet to analyze clinical notes from the first 72 hours of admission. Model outputs were compared against a physician-adjudicated gold standard in a cohort of 1,000 encounters from patients with cirrhosis who presented to the ED. Performance was benchmarked against ICD-10 code-based labeling and CDC Adult Sepsis Event criteria. ResultsINFEHR achieved 94.7% overall accuracy, with 99.5% sensitivity and 92.8% positive predictive value for identifying infection presence, outperforming ICD-10-based classification across all metrics (p < 0.0001). The model also demonstrated strong performance in classifying pathogen type and infection site. This pipeline processed notes within seconds, offering improvements in efficiency and scalability over manual review. ConclusionINFEHR offers a scalable, reproducible, and accurate method for infection phenotyping in cirrhosis. By overcoming limitations of traditional coding and manual review, it supports high-throughput infection surveillance, improves cohort construction for clinical research, and enables future integration into real-time decision-support tools in hepatology.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A deep learning model for clinical outcome prediction using longitudinal inpatient electronic health records 93%
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 90%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 90%
Similar papers in this journal
- Deep representation learning for clustering longitudinal survival data from electronic health records 92%
- Integration of clinical characteristics, lab tests and a deep learning CT scan analysis to predict severity of hospitalized COVID-19 patients 91%
- Integrating a host transcriptomic biomarker with a large language model for diagnosis of lower respiratory tract infection 91%
Similar papers in this journal
- Community-acquired pneumonia identification from electronic health records in the absence of a gold standard: a Bayesian latent class analysis 94%
- Can the application of machine learning to electronic health records guide antibiotic prescribing decisions for suspected urinary tract infection in the Emergency Department? 91%
- Hospital-wide Natural Language Processing summarising the health data of 1 million patients 91%
Similar papers in this journal
- Characterizing Long COVID: Deep Phenotype of a Complex Condition 92%
- irAE-GPT: Leveraging large language models to identify immune-related adverse events in electronic health records and clinical trial datasets 91%
- Transformer-based deep learning model for the diagnosis of suspected lung cancer in primary care based on electronic health record data 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.