A single-patient task exposes a failure of safety alignment in clinical language models
Gorenshtein, A.; Jia, E. L.; Omar, M.; Brook, O. R.; Ahmed, M.; Kruskel, J. B.; Barash, Y.; Klang, E.
Show abstract
Safety alignment should persist while a language model performs a task. We tested whether a single-patient triage task suppressed a warning about a second patient. Each case centered on Patient 1; Patient 2's urgent problem appeared only in passing. Sixteen models saw each case twice: once as a general assistant and once while producing a triage record for Patient 1. As general assistants, models warned the caller in 87% of cases; under the task, they did so in 21%. Every model showed a significant decrease. Yet under the task, the record still mentioned Patient 2 in 76% of cases and recommended urgent care in 67%. Across 15 open-weight models, repeating the emergency-care instruction raised the warning rate only to 29%; moving the message-to-caller field to the top raised it to 36%. Current safety alignment did not reliably persist under task assignment.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 95%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 93%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 92%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A Study of Calibration as a Measurement of Trustworthiness of Large Language Models in Biomedical Research 92%
- Evaluation of Patient-Level Retrieval from Electronic Health Record Data for a Cohort Discovery Task 92%
- Natural Language Processing for Automated Annotation of Medication Mentions in Primary Care Visit Conversations 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.