PAUSE-Agents: A Clinician-in-the-Loop Multi-Agent AI Pipeline for ICU-to-Ward Handoff Briefs
Amagai, S.; Liao, W.-T.; Murphy, C.; Reamer, C.; Liu, Y.; Ambil, B.; Fernandes, G.; Santhosh, L.; Lyons, P.; Jordan, N.; Liebovitz, D.; Kline, A.; Rojas, J. C.; Luo, Y.; Gao, C.
Show abstract
ICU-to-ward transfers are high-risk transitions marked by information loss and burdensome handoff preparation. We developed PAUSE-Agents, a clinician-in-the-loop multi-agent LLM pipeline that drafts source-attributed handoff briefs from structured ICU data and clinical notes using the clinician-developed ICU-PAUSE template. Mirroring ICU team structure, PAUSE-Agents routes each record through a scribe extractor, 6 role-specialized agents, explicit conflict surfacing, and deterministic safety checks before synthesis, producing an editable first draft rather than an autonomous note. In a single-center medical ICU cohort, 5 physicians completed 100 reviews of 84 agent-drafted briefs. Among adjudicable claims, 98.8% were verified and 1.2% were incorrect; 88% of briefs had no pertinent omission, and mean PDSQI-9 quality was 4.20/5. PAUSE-Agents surfaced 118 conflict warnings and 421 safety flags, making documentation inconsistencies visible before handoff. An o4-mini PDSQI-9 judge showed limited case-level discrimination but supported aggregate monitoring. We release PAUSE-Agents and its clinician evaluation application.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 91%
Similar papers in this journal
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 92%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 91%
- Understanding COVID-19 trajectories from a nationwide linked electronic health record cohort of 56 million people: phenotypes, severity, waves & vaccination 91%
Similar papers in this journal
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 93%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 92%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 92%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 92%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 89%
- A Crowdsourcing Approach to Develop Machine Learning Models to Quantify Radiographic Joint Damage in Rheumatoid Arthritis 88%
Similar papers in this journal
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
- Evaluation of Domain Generalization and Adaptation on Improving Model Robustness to Temporal Dataset Shift in Clinical Medicine 93%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.