MTS-Bench: A Manchester Triage System Benchmark for Language Model Triage Safety
Ravichandran, S.; Romano, M.; Corga da Silva, R.; Mendes, T.; Absi, N.; Isidoro, M.; Kumar, S.; Van der Heijden, M.; Gnanapragasam, V. E.
Show abstract
Background: General purpose language models such as ChatGPT are increasingly used by physicians and triage nurses during emergency triage. A recent study reported 51.6% undertriage of emergencies when patients queried ChatGPT directly (Ramaswamy et al., 2026). DR. INFO is an agentic AI based clinical assistant that retrieves over a curated clinical knowledge base, and an MTS specific retrieval configuration is available in which the system also retrieves the Manchester Triage System (MTS) textbook at inference time. The safety of these systems as a triage adjunct against a structured framework has not been characterised. Methods: We adapted the clinical scenarios published by Ramaswamy et al. and mapped them to the Manchester Triage System, yielding 39 emergency cases covering all five MTS priority levels. Each case was evaluated in two variants, one without and one with the objective clinical data block (vital signs, examination findings, and laboratory results), and permuted across two genders, giving 156 prompts per condition. Three systems were tested with and without a misleading GP referral statement prepended as an anchoring statement, giving 312 prompts per system: DR. INFO Baseline, DR. INFO with MTS retrieval, and OpenAI GPT-5.1. The primary outcome was the undertriage rate on the ordered MTS scale, tested with Fisher's exact test. Results: GPT-5.1 undertriaged 44.2% of cases (69/156; 95% CI 36.7 to 52.1), including 75.0% of Red and 73.4% of Orange presentations. Both DR. INFO configurations undertriaged 11.5% of cases (18/156; 95% CI 7.4 to 17.5; Fisher's exact p = 1.0 x 10^-10 versus GPT-5.1). GPT-5.1 produced 6 dangerous misses (3.8%), and both DR. INFO configurations produced none (p = 0.030). When the anchoring statement was prepended, GPT-5.1 undertriaged 8 of 8 Red cases, while both DR. INFO configurations continued to undertriage none. Adding objective clinical data to the input reduced undertriage in DR. INFO with MTS retrieval from 19.2% to 3.8% (p = 0.005). DR. INFO Baseline and GPT-5.1 showed no comparable change. There was no significant effect of gender. Conclusion: On this benchmark, replacing a general purpose language model with an agentic retrieval augmented system over a curated clinical knowledge base substantially reduced the undertriage and dangerous miss rates. Adding retrieval of the Manchester Triage System textbook to the agentic system was further associated with a reduced susceptibility to the anchoring statement and with an appropriate change in the assigned MTS priority when objective clinical data became available. Of the three configurations evaluated here, only DR. INFO with MTS retrieval combined a clinically conservative assignment at first contact with appropriate updating as additional clinical information arrived.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Emergency medicine patient wait time multivariable prediction models: a multicentre derivation and validation study 91%
- Accuracy of the National Early Warning Score version 2 (NEWS2) in predicting need for time-critical treatment: Retrospective observational cohort study 89%
- Evaluating the impact of a pulse oximetry remote monitoring programme on mortality and healthcare utilisation in patients with covid-19 assessed in Accident and Emergency departments in England: a retrospective matched cohort study 89%
Similar papers in this journal
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 94%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- Comparison of local large language models for extraction of signs and symptoms data from electronic health records 91%
Similar papers in this journal
- Measuring the Quality of AI-Generated Clinical Notes: A Systematic Review and Experimental Benchmark of Evaluation Methods 92%
- The role of natural language processing in cancer care: a systematic scoping review with narrative synthesis 91%
- Building Large-Scale Registries from Unstructured Clinical Notes using a Low-Resource Natural Language Processing Pipeline 90%
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 94%
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.