Beyond Accuracy: An Efficiency- and Safety-Aware Framework for Evaluating Clinical AI with Large Language Models
Zaki, N.; Akor, A.; Aburuz, S.; ZainAlAbdin, S.
Show abstract
BackgroundLarge language models (LLMs) demonstrate strong performance on medical reasoning tasks, but current evaluation approaches focus primarily on accuracy, neglecting the efficiency-safety trade-offs critical for real-world clinical utility. MethodsWe developed and validated the Clinical Value Density (CVD) framework, a novel metric quantifying clinical utility per unit of cognitive resource consumed. Six state-of-the-art LLMs (GPT-4o, Gemini-2.5, Claude-Sonnet-4, Grok-3, DeepSeek-R1, and Kimi-K2) were evaluated across 60 authentic clinical pharmacology scenarios derived from UAE healthcare practice and benchmarked against board-certified clinical pharmacist responses. Performance was assessed across efficiency, semantic similarity, safety, relevance, consistency, and conciseness, with triangulated validation against clinician preference and task efficiency. ResultsTraditional metrics (BLEU: 0.003-0.024; ROUGE-L: 0.079-0.166) failed to reflect clinical utility, while the CVD framework exposed critical efficiency-safety trade-offs. GPT-4o achieved the highest normalized CVD (0.475), delivering fourfold efficiency gains over pharmacists (41 vs 178 tokens) but with moderate safety (0.437), requiring supervised deployment. Grok-3 and Gemini-2.5 led in safety (0.605 and 0.542) but at the cost of efficiency. DeepSeek-R1 and Kimi-K2 produced unsafe brevity-accuracy trade-offs, generating concise but clinically inaccurate responses. Comparative validation revealed strong alignment between CVD and task efficiency (r = 0.924), but divergence from clinician preference (r = -0.845), reflecting a bias toward verbose outputs. ConclusionsCurrent LLMs cannot reliably perform clinical functions autonomously. Instead, they are best positioned for AI-assisted, supervised integration, where efficiency and safety are balanced under professional oversight. The CVD framework provides a potentially useful for regulatory evaluation for quantifying deployment readiness, aligning AI evaluation with the real-world constraints of cognitive load, time, and patient safety. Future research should extend CVD across specialties, scale validation datasets, and conduct real-time workflow trials to establish specialty-specific safety thresholds for eventual autonomy. Author SummaryLarge language models (LLMs) are impressively accurate on medical reasoning benchmarks, yet clinical work is not a quiz, it is a race against time under strict safety constraints. We introduce Clinical Value Density (CVD), a simple yet powerful metric that captures clinical utility per unit of cognitive cost (i.e., how much useful care an answer delivers for the effort it imposes on clinicians). In a head-to-head evaluation of six state-of-the-art LLMs across 60 authentic clinical pharmacology scenarios benchmarked against board-certified pharmacists, traditional metrics (BLEU/ROUGE) barely moved, while CVD exposed the true efficiency-safety trade-offs that determine bedside usefulness. For example, GPT-4o delivered the highest normalized CVD (0.475) with [~]4x fewer tokens than pharmacists ({approx}41 vs 178), but only moderate safety (0.437) appropriate for supervised use, not autonomy. Conversely, Grok-3 and Gemini-2.5 were safest but less efficient; DeepSeek-R1 and Kimi-K2 risked unsafe brevity- accuracy trade-offs. Crucially, CVD strongly tracked task efficiency (r=0.924) yet diverged from clinician preference (r=-0.845), revealing a bias toward overly verbose answers that feel reassuring but slow care.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 95%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- Dataset Documentation for Responsible AI: Analysis of Suitability and Usage for Health Datasets 94%
Similar papers in this journal
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 94%
- COHD-COVID: Columbia Open Health Data for COVID-19 Research 94%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 94%
Similar papers in this journal
- Artemis: Harnessing Knowledge Graphs for Next-Generation Drug Target Prioritization 90%
- Network-based estimation of therapeutic efficacy and adverse reaction potential for prioritisation of anti-cancer drug combinations 90%
- Deep Learning-based Framework for Mycobacterium tuberculosis Bacterial Growth Detection for Antimicrobial Susceptibility Testing 89%
Similar papers in this journal
- Developing A Deep Learning Natural Language Processing Algorithm For Automated Reporting Of Adverse Drug Reactions 94%
- Biomedical Text Normalization through Generative Modeling 93%
- Demonstrating the Consequences of Learning Missingness Patterns in Early Warning Systems for Preventative Health Care: A Novel Simulation and Solution 93%
Similar papers in this journal
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 97%
- Usability of a Machine-Learning Clinical Order Recommender System Interface for Clinical Decision Support and Physician Workflow 95%
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.