Are Frontier Large Language Models Safer Than Government-Backed Symptom Checkers for Clinical Self-Triage? A Standardised Vignette Evaluation
Chowdhury, A. R.; Chowdhury, B.
Show abstract
Background: Consumer use of AI chatbots for health advice is rising, yet triage safety relative to established services remains unclear. Australia's Healthdirect, a government-backed symptom checker with 2.4 million uses in FY2024-25, remains unevaluated against frontier large language models (LLMs), and whether premium subscriptions improve triage safety remains unexplored. This study compared the triage accuracy and safety of Healthdirect against six LLM configurations across ChatGPT, Claude, and Gemini, assessed whether paid subscriptions improve triage safety, and characterised each system's error patterns. Methods: Forty-five clinical vignettes from the Semigran et al. benchmark spanning emergency, non-emergent, and self-care categories (15 each) were evaluated across seven systems. Healthdirect was tested following a seven-rule interaction protocol. LLMs were evaluated using first-person patient-language prompts under free-tier and paid-tier conditions. Outcomes were triage accuracy, emergency sensitivity, under-triage, and critical misses, analysed using Cochran's Q, Bonferroni-corrected McNemar tests, Cohen's kappa, and Wilson intervals. Findings: Triage accuracy differed significantly (Cochran's Q = 36.79, p < 0.001). Healthdirect achieved 48.9% accuracy (95% CI 35.0% to 63.0%; kappa = 0.233) versus 73.3% to 86.7% for LLMs (kappa = 0.600 to 0.800). Healthdirect operated under conservative interactive defaults while LLMs received complete information in a single prompt, which may have disadvantaged Healthdirect. Emergency sensitivity was 46.7% versus 80.0% to 86.7% for LLMs. Healthdirect produced two critical misses; no LLM produced any across 270 evaluations (95% CI 0% to 1.4%). When LLMs undertriaged, they recommended GP care rather than self-care. No tier differences were significant (all p > 0.05), and most systems over-triaged self-care cases. Interpretation: Frontier LLMs demonstrated higher triage accuracy and safer error profiles than Healthdirect. All LLMs avoided critical misses; Healthdirect did not. Premium subscriptions did not significantly improve triage safety. These findings support clinical governance decisions about whether LLMs warrant formal evaluation alongside government-backed symptom checkers.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Impact of the Federated Data Platform's digital surgery scheduling system on elective theatre utilisation at an NHS Trust: an interrupted time series analysis 93%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 93%
Similar papers in this journal
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
- Benefits and Challenges of Using Virtual Primary Care During the COVID-19 Pandemic: From Key Lessons to a Framework for Implementation 92%
- From Patient Voices to Policy: Data Analytics Reveals Patterns in Ontarios Hospital Feedback 92%
Similar papers in this journal
- What Do Clinicians Edit in Ambient AI-Drafted Clinical Documentation? A Qualitative Content Analysis 94%
- Measure what matters: counts of hospitalized patients are a better metric for health system capacity planning for a reopening 93%
- Empowering Personalized Pharmacogenomics with Generative AI Solutions 92%
Similar papers in this journal
- Feasibility trial of a new digital training package to enhance primary care practitioners' communication of clinical empathy and realistic optimism 93%
- Triaging and Referring In Adjacent General and Emergency Departments (the TRIAGE trial): a cluster randomised controlled trial 93%
- Essential Indicators of Quality in Primary Care Settings: An Evidence-Based, Structured, Expert Approach 93%
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 94%
- Ambient scribe in general practice: a multi-perspective before-after longitudinal mixed-methods study 93%
- Understanding digital health technology implementation in rehabilitation: Development of the Rehabilitation Technologies Implementation model 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.