Aggregate benchmark scores obscure patient safety implications of errors across frontier language models
Linzmayer, R.; Ramaswamy, A.; Hugo, H.; Nadkarni, G.; Elhadad, N.
Show abstract
Frontier language models are widely used for health-related queries, yet aggregate benchmark scores do not capture safety implications of errors. We applied the recent Nature Medicine triage benchmark across nine frontier models, comparing directional error profiles, contextual bias, and crisis calibration. In-range accuracy ranged from 75.0% to 87.7%, obscuring clinically meaningful error differences. Looking at directionality of errors, under-triage ranged from 0.0% (GPT-5.2) to 12.3% (GPT-5-mini), over-triage varied independently (9.4-36.9%), and under-triage was uncorrelated with aggregate accuracy. When family members minimized symptoms, all models tested shifted toward lower acuity in ambiguous cases (OR range 2.9-14.9), the only contextual effect observed consistently, and access barriers increased under-triage risk in six. Suicide crisis resource mention rates were low and variable across all models. This cross-model heterogeneity and non-monotonic performance across model generations show that aggregate accuracy alone cannot characterize, rank, or predict the clinical safety of deployed language models.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Analysis of Eligibility Criteria Clusters Based on Large Language Models for Clinical Trial Design 92%
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 92%
- Large Language Models Facilitate the Generation of Electronic Health Record Phenotyping Algorithms 92%
Similar papers in this journal
- Optimal policy determination in sequential systemic and locoregional therapy of oropharyngeal squamous carcinomas: A patient-physician digital twin dyad with deep Q-learning for treatment selection 90%
- Structured Codes and Free-Text Notes: Measuring Information Complementarity in Electronic Health Records 90%
- Virtual health care for community management of patients with COVID-19. 90%
Similar papers in this journal
Similar papers in this journal
- Modular Clinical Decision Support Networks (MoDN)—Updatable, Interpretable, and Portable Predictions for Evolving Clinical Environments 92%
- Hospital-wide Natural Language Processing summarising the health data of 1 million patients 90%
- Automated identification of abnormal infant movements from smart phone videos 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.