Back

Automating the triage of rheumatology outpatient referrals: a comparative evaluation of 23 large language models under simple and advanced prompting

Roberts, L.

2026-08-10 health systems and quality improvement
10.64898/2026.08.05.26359488 medRxiv
Show abstract

Objective. Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify to optimal approach. Methods. Twenty referral scenarios spanning the urgency spectrum, based on real referrals were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded, to produce a consensus reference standard. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per condition). The experiment was run with a simple prompt and repeated with a advanced prompt supplying explicit triage expectations and worked examples. Results. All 2760 attempts returned valid categories. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho=0.42; P=.047) and accuracy tracked cost. Advanced prompting minimised between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P=.01), abolished the size-accuracy association (rho=-0.05; P=.83) and removed the accuracy-cost relationship. Leading models matched expert consensus on most cases, within or above the range reported for human triage. Under-triage errors persisted with some LLMs. Conclusion. Contemporary LLMs categorise rheumatology referral urgency as well or better than published human triage systems. Advanced LLM prompting methods substitute for the reasoning capability of larger models, suggesting that LLM performance on this task may not require the most expensive models. The tools to automate this administrative task appear to already exist. Strong candidate LLMs that might serve a production ready solution have been identified.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.