Benchmarking Large Language Models for Intensive Care Unit Clinical Decision Support: A Dual Safety Evaluation of 26 Models on Consumer Hardware
Shlyakhta, T.
Show abstract
BackgroundLarge Language Models (LLMs) show promise for clinical decision support in Intensive Care Units (ICU), but their safety and reliability remain inadequately evaluated through dual testing of both memory-dependent and memory-independent safety mechanisms. ObjectiveTo comprehensively evaluate LLMs using two independent safety tests: context-dependent contraindication memory (penicillin allergy recall) and context-independent authority resistance (Extended Milgram Test), revealing whether these represent unified or dissociated safety mechanisms. MethodsTwenty-three LLMs underwent automated testing via 24-hour ICU simulation on consumer hardware (NVIDIA RTX 3060 12GB). A subset of 26 models completed an Extended Milgram Test with five escalating harmful command scenarios. Scoring assessed safety compliance, Milgram resistance, conflict detection, and performance. ResultsCritical findings revealed dissociation between abstract ethics and clinical memory. While 65% of models achieved perfect Milgram resistance (100%), only 8.7% (n=2) correctly refused penicillin with allergy mention. Eight models demonstrated 100% Milgram resistance yet failed allergy recall (r = -0.39, p = 0.23). Only Granite 3.1 8B achieved perfect performance on both tests. ConclusionsAbstract ethical reasoning (refusing harmful orders in principle) is independent from concrete clinical memory (tracking patient-specific risks). Safe medical AI requires both capabilities--rarely both present. Dual safety testing should become mandatory for medical AI certification. HighlightsO_LIOnly 8.7% of tested LLMs passed critical safety tests for medication prescribing C_LIO_LIFirst study demonstrating dissociation between abstract ethics and clinical memory (r = -0.39) C_LIO_LIEight models refused all harmful orders but forgot documented allergies C_LIO_LIGranite 3.1 8B only model achieving perfect performance on both safety tests C_LIO_LIDual safety testing framework proposed for medical AI certification C_LI
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and Prospective Implementation of a Large Language Model based System for Early Sepsis Prediction 95%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 94%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 94%
Similar papers in this journal
- Raising awareness of potential biases in medical machine learning: Experience from a Datathon 94%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 94%
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 94%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 93%
- Validation of a Derived International Patient Severity Algorithm to Support COVID-19 Analytics from Electronic Health Record Data 92%
Similar papers in this journal
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 91%
- Development of a customised data management system for a COVID-19-adapted colorectal cancer pathway 90%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 90%
Similar papers in this journal
- Development and Validation of ‘Patient Optimizer’ (POP) Algorithms for Predicting Surgical Risk with Machine Learning 92%
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 92%
- Implicit bias in Critical Care Data: Factors affecting sampling frequencies and missingness patterns of clinical and biological variables in ICU Patients 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.