Leveraging simulation to provide a practical framework for assessing the novel scope of risk of LLMs in healthcare
Kalinich, M.; Luccarelli, J.; Moss, F.; Torous, J.
Show abstract
Structured AbstractO_ST_ABSBackgroundC_ST_ABSLarge language models (LLMs) are rapidly entering clinical care, yet their definitionally probabilistic outputs have delivered a variety of grossly unsafe responses to users. The difficulty in quantifying and mitigating the novel risks posed by LLMs threatens to stall the regulatory evaluation and clinical deployment of LLM-based software as a medical device (LLM-SaMD). A practical, evidence-based framework is urgently needed for extending existing medical-device regulations to encompass LLM-SaMDs. Using synthetic interactions between a chatbot and a potentially suicidal user, we demonstrate a simulation-based framework that provides a reproducible and generalizable method for evaluating the novel risks of LLM-SaMDs. MethodsWe developed a framework integrating LLM performance testing into SaMD risk estimation. Fourteen open-source models ranging from 270 million to 70 billion parameters (Qwen, Gemma, and LLaMA families) were evaluated on three safety-classification tasks: suicidal-ideation detection, therapy-request detection, and therapy-like interaction detection. Synthetic datasets were generated by Gemini 2.5 Pro and verified by psychiatrists. Model false-negative rates informed probabilistic estimates of P1, the likelihood of a hazard progressing to a hazardous situation, and P2, the likelihood of that situation resulting in harm. ResultsLLM success at generating synthetic safety datasets varied substantially by task, with strong performance for neutral and non-therapeutic content but frequent errors in suicidal-ideation and therapy-like interactions. Across 14 models (270 million-70 billion parameters), performance generally improved with size but included notable outliers. Estimated P1 values (hazard to hazardous situation) ranged from 2.0x10-8 to 2.6x10-4 and P2 (hazardous situation to harm) from 7.1x10-5 to 9.6x10-3, spanning up to four orders of magnitude. ConclusionSimulation extends existing device-safety frameworks to address the novel risks of large language models. Rather than replacing regulatory judgment, it provides a reproducible method for quantifying uncertainty, clarifying assumptions, and linking model failures to plausible harms. Our case example demonstrates a generalizable approach that can overcome current regulatory barriers while remaining practical for manufacturers and regulators, supporting timely and transparent oversight that keeps patients safe while avoiding unnecessary barriers to delivering the clinical promise of LLM-based medical devices. Brief DescriptionThis study introduces a quantitative framework for evaluating and mitigating the unique risks that large language models (LLMs) pose in healthcare. By mapping the pathways from LLM-generated hazards to harms onto existing regulatory risk-analysis structures and estimating the probability of these transitions through computational simulation, the framework empirically bounds uncertainty and identifies where real-world evidence is needed to validate and monitor model performance before, during, and after clinical deployment.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 93%
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 93%
- Toward Trustworthy Chatbots: A Protocol for Red Teaming for Health Related Conversations 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.