Agentic Chart Review from Longitudinal Clinical Notes: a Lung Cancer Guideline Concordance Use Case
Jiang, Y.; He, X.; Ai, X.; Jalal, S.; Maniar, R.; Majji, R. K.; Zhang, Y.; Liu, J.; Fedele, D.; Zhuang, Y.; Hollenbach, J.; Bian, J.
Show abstract
Clinical chart abstraction extracts structured patient variables from longitudinal clinical notes but is labor-intensive and difficult to scale. We evaluated LLM agents for question-guided chart review using lung cancer molecular testing guideline concordance as a use case. Two configurations were compared: (1) sequential note review using metadata and chronology, and (2) the same framework augmented with keyword-based note search. Gold-standard labels were established by human annotators. The search-enabled agent achieved higher accuracy (92.4% vs. 83.5%) and reduced errors by more than half (41 vs. 89) by retrieving evidence from long, heterogeneous note histories. In guideline concordance evaluation, most determinate patient-rule assessments were concordant (80.7%), while most apparent non-concordance reflected missing molecular testing documentation rather than documented care deviations. These results suggest tool-augmented LLM agents can approximate key aspects of human chart review and support scalable information extraction from longitudinal clinical documentation.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 93%
Similar papers in this journal
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 93%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 90%
- Missing data in the medical record for oncology patients: prevalence and association with outcomes 90%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.