Automating the triage of rheumatology outpatient referrals: a comparative evaluation of 23 large language models under simple and advanced prompting
Roberts, L.
Show abstract
Objective. Triage of rheumatology outpatient referrals is a high-volume administrative task that consumes senior specialist time without advancing patient care. The human triage system is only moderately accurate and reproducible. We assessed whether contemporary large language models (LLMs) are able to perform well enough to support automating this task in practice. In addition, the effects of different prompting techniques on triage accuracy and cost was assessed to help identify to optimal approach. Methods. Twenty referral scenarios spanning the urgency spectrum, based on real referrals were created by a certified Australian rheumatologist. Four rheumatologists triaged all cases independently and blinded, to produce a consensus reference standard. Twenty-three LLMs each triaged every referral into one of five urgency categories, three times (1380 outputs per condition). The experiment was run with a simple prompt and repeated with a advanced prompt supplying explicit triage expectations and worked examples. Results. All 2760 attempts returned valid categories. Under the simple prompt, performance separated into distinct tiers, larger models were more accurate (Spearman rho=0.42; P=.047) and accuracy tracked cost. Advanced prompting minimised between-model variance in accuracy 5.3-fold (0.014 to 0.003; Levene P=.01), abolished the size-accuracy association (rho=-0.05; P=.83) and removed the accuracy-cost relationship. Leading models matched expert consensus on most cases, within or above the range reported for human triage. Under-triage errors persisted with some LLMs. Conclusion. Contemporary LLMs categorise rheumatology referral urgency as well or better than published human triage systems. Advanced LLM prompting methods substitute for the reasoning capability of larger models, suggesting that LLM performance on this task may not require the most expensive models. The tools to automate this administrative task appear to already exist. Strong candidate LLMs that might serve a production ready solution have been identified.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Large Language Models in Real-World Clinical Workflows: A Systematic Review of Applications and Implementation 90%
- Human-supervised, large language model-based clinical decision support aligned to national newborn protocols in Kenya: a pragmatic, early-stage evaluation 89%
- Medical Clinical Minds Meet Artificial Intelligence: Italian Physicians' Knowledge, Attitudes, and Concordance between Italian Physicians and AI-Generated Diagnoses. A National Cross-Sectional Study 88%
Similar papers in this journal
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- A method for rapid machine learning development for data mining with Doctor-In-The-Loop 90%
- Refractory Inflammatory Arthritis definition and model generated through patient and multi-disciplinary professional modified Delphi process 90%
Similar papers in this journal
- Spatio-temporal modelling of referrals to outpatient respiratory clinics in the integrated care system of the Morecambe Bay area, England 89%
- Revisiting the use and effectiveness of patient-held records in rural Malawi 88%
- Modelling vaccination capacity at mass vaccination hubs and general practice clinics 88%
Similar papers in this journal
- Discrepancy review: A feasibility study of a novel peer review intervention to reduce undisclosed discrepancies between registrations and publications 86%
- Transparency in the secondary use of health data: Assessing the status quo of guidance and best practices 86%
- Communicating personalised risks from COVID-19: guidelines from an empirical study 85%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.