LLMs Can Do Medical Harm: Stress-Testing Clinical Decisions Under Social Pressure
Omar, M.; Agbareia, R.; McGreevy, J.; Gorenshtein, A.; Charney, A.; Sakhuja, A.; Glicksberg, B. S.; Nadkarni, G.; Klang, E.
Show abstract
BackgroundLarge language models (LLMs) are entering clinical workflows, yet their effect on clinical decisions and potential for harm are uncertain. MethodsWe measured harmful decision output from an ensemble of 20 LLMs across >10 million clinical scenarios with safety or ethical dilemmas. Each case was shown under a neutral control and six Milgram-style social-pressure conditions, with or without a brief mitigation cue ("verify or escalate if unsafe"). The primary outcome was the proportion of potentially harmful responses. We used two-proportion tests/{chi}2 tests and confirmatory mixed-effects logistic models. ResultsAcross all runs (N = 10,096,800), LLMs produced 1.18 million potentially harmful outputs (11.7%). Mitigation reduced harmful decisions from 16.6% to 10.1% (p < 0.001). When exposed to social pressure, models behaved predictably but unevenly: prompts framed as authority or responsibility transfer generated the most harmful responses, whereas control prompts, neutral and pressure-free, produced the fewest (mitigated 8.3-9.6%; unmitigated 14.3-16.0%; {chi}2 p < 0.001). In other words, when told what to do, or told that someone else would take responsibility, models were more likely to comply, even when the instruction was unsafe. These effects were consistent across datasets and models ConclusionLLMs can generate harmful medical decisions at scale. A brief safety reminder reduces, but does not eliminate, this behavior. These results highlight the need to measure harm propensity as a core performance metric and to maintain guardrails and continuous physician oversight before integrating LLMs into clinical decision-making.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 93%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 93%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 93%
Similar papers in this journal
- Generalizability Challenges of Mortality Risk Prediction Models: A Retrospective Analysis on a Multi-center Database 93%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 92%
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 92%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 94%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 91%
- Clinical Outcomes, Costs, and Cost-effectiveness of Strategies for People Experiencing Sheltered Homelessness During the COVID-19 Pandemic 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.