Clinical Agents Don't Care
Klang, E.; Glicksberg, B. S.; Gorenshtein, A.; Gavin, N.; Freeman, R.; Stump, L.; Charney, A. W.; Ting, D. S. W.; Omar, M.; Nadkarni, G.
Show abstract
BackgroundLarge language models (LLMs) now power clinical agents that can plan, call tools, and write into electronic health records (EHRs). They are becoming actors, not assistants. Given known LLM faults, quality assurance is essential before clinical use. A key question is whether agents notice patient-identity errors or act indifferent. MethodsWe created a record environment using publicly available MIMIC-IV real-world emergency department data. Agents were instructed to copy ICD-10 codes from visit headers into patient records using Extract and Store tools, with an option to record "UNKNOWN" if uncertain or abstain. Each agent was presented with ten batched records from the same patient (clean version). Then we tampered one of the records and evaluated how the agent responded. We ran four separate batches: the clean baseline batch, a batch with one visit with a complete swapped header from another patient, a batch with one visit with a one-digit MRN change, and a batch with age shifted in one visit. Six models, both closed- and open-weight, completed 1.2 million tool calls to assess model performance. The endpoint was whether agents would identify when fields were inconsistent identity. ResultsAgents frequently failed, copying codes into tampered charts. GPT-4.1 flagged mismatched headers as UNKNOWN in 17.4% of runs but never detected subtle faults. GPT-4.1-nano detected 4.4% of header swaps and <1% of MRN or age changes. GPT-5-chat never identified mismatches but omitted responses in 12.6% of cases. Other models rarely abstained. Subtle tampering passed almost entirely without detection. ConclusionsClinical agents are often indifferent to patient details inconsistencies. The central risk is misbinding, not miscoding. Safe deployment requires explicit identity verification, abstention when uncertain, and benchmarks that treat record integrity, not just accuracy, as a primary outcome.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale 94%
- MMFP-Tableau: Enabling Precision Mitochondrial Medicine through Integration, Visualization, and Analytics of Clinical and Research Health System Electronic Data 92%
- Understanding Data Differences across the ENACT Federated Research Network 92%
Similar papers in this journal
Similar papers in this journal
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 92%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 91%
- Distinguishing Admissions Specifically for COVID-19 from Incidental SARS-CoV-2 Admissions: A National Retrospective EHR Study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.