Personas Shift Clinical Action Thresholds in Large Language Models
Klang, E.; Gorenshtein, A.; Omar, M.; Nadkarni, G.
Show abstract
Background and aimsClinical LLM deployment is shifting from feasibility to liability, while current guidance largely treats model behavior as a control problem. We tested whether decision-style system prompts shift clinical action thresholds when clinical facts are held constant, and whether these shifts are consistent across settings and models. MethodsWe defined nine physician personas by crossing three ethical orientations (duty-, care-, utilitarian) with three cognitive styles (intuitive, integrative, analytic). Twenty open-weight LLMs were evaluated on 2,500 simulated ED vignettes and 2,500 MIMIC-IV-Note discharge summaries. For each text, models answered five binary decision items (safety, autonomy, treatment, resource use, follow-up). Each condition was repeated ten times, yielding 5,000,000 total decisions. ResultsUnder baseline prompting, models answered "Yes" to 42.8% of decisions. Persona prompts shifted affirmative rates from 36.9% to 46.4%, a 9.5-percentage-point swing under fixed clinical evidence. Effects were largest in autonomy and treatment and were consistent across corpora (85.7% directional agreement; r = 0.82 for effect sizes). Susceptibility varied by model (4.9-16.1 points), with no consistent protection from medical fine-tuning or model size. ConclusionsDecision-style system prompts reliably change clinical action thresholds in LLMs under fixed evidence. Prompting is a policy-setting layer, not just a communication layer, and should be treated as a first-class deployment configuration.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 93%
Similar papers in this journal
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 94%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 94%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 91%
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 91%
Similar papers in this journal
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 92%
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 91%
- Predicting hospital-onset COVID-19 infections using dynamic networks of patient contacts: an observational study 91%
Similar papers in this journal
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 93%
- Cracking the Code: A Scoping Review to Unite Disciplines in Tackling Legal Issues in Health Artificial Intelligence 90%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.