An agentic AI system enhances clinical detection of immunotherapy toxicities: a multi-phase validation study
Gallifant, J.; Chen, S.; Shin, K.-Y.; Kellogg, K. C.; Doyle, P. F.; Guo, J.; Ye, B.; Warrington, A.; Zhai, B. K.; Hadfield, M. J.; Gusev, A.; Ricciuti, B.; Christiani, D. C.; Aerts, H. J.; Kann, B. H.; Mak, R. H.; Nelson, T. L.; Nguyen, P.; Schoenfeld, J. D.; Topaloglu, U.; Catalano, P.; Hochheiser, H. H.; Warner, J. L.; Sharon, E.; Kozono, D. E.; Savova, G. K.; Bitterman, D.
Show abstract
Immune-related adverse events (irAEs) affect up to 40% of patients receiving immune checkpoint inhibitors, yet their identification depends on laborious and inconsistent manual chart review. Here we developed and evaluated an agentic large language model system to extract the presence, temporality, severity grade, attribution, and certainty of six irAE types from clinical notes. Retrospectively (263 notes), the system achieved macro-averaged F1 of 0.92 for detection and 0.66 for multi-class severity grading; self-consistency improved F1 by 0.14. The best-performing configuration cost approximately $0.02 per note. In prospective silent deployment over three months (884 notes), detection F1 was 0.72-0.79. In a randomized crossover study of clinical trial staff (17 participants, 316 observations), agentic assistance reduced annotation time by 40% (P < 0.001), increased complete-match accuracy (OR 1.45; 95% CI 1.01-2.09; P = 0.045), and improved inter-annotator agreement (Krippendorffs from 0.22-0.51 to 0.82-0.85). These results demonstrate that agentic AI coupled with human verification could enhance efficiency, performance, and consistency for irAE assessment.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 96%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 93%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 94%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 91%
- Missing data in the medical record for oncology patients: prevalence and association with outcomes 89%
Similar papers in this journal
Similar papers in this journal
- Systematic Review of Large Language Models for Patient Care: Current Applications and Challenges 94%
- Achieving Inclusive Healthcare through Integrating Education and Research with AI and Personalized Curricula 93%
- Pretrained Patient Trajectories for Adverse Drug Event Prediction Using Common Data Model-based Electronic Health Records 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.