Accuracy of Online Symptom-Assessment Applications, Large Language Models, and Laypeople for Self-Triage Decisions: A Systematic Review
Kopka, M.; von Kalckreuth, N.; Feufel, M. A.
Show abstract
Symptom-Assessment Application (SAAs, e.g., NHS 111 online) that assist medical laypeople in deciding if and where to seek care (self-triage) are gaining popularity and their accuracy has been examined in numerous studies. With the public release of Large Language Models (LLMs, e.g., ChatGPT), their use in such decision-making processes is growing as well. However, there is currently no comprehensive evidence synthesis for LLMs, and no review has contextualized the accuracy of SAAs and LLMs relative to the accuracy of their users. Thus, this systematic review evaluates the self-triage accuracy of both SAAs and LLMs and compares them to the accuracy of medical laypeople. A total of 1549 studies were screened, with 19 included in the final analysis. The self-triage accuracy of SAAs was found to be moderate but highly variable (11.5 - 90.0%), while the accuracy of LLMs (57.8 - 76.0%) and laypeople (47.3 - 62.4%) was moderate with low variability. Despite some published recommendations to standardize evaluation methodologies, there remains considerable heterogeneity among studies. The use of SAAs should not be universally recommended or discouraged; rather, their utility should be assessed based on the specific use case and tool under consideration.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 95%
- How and why do Quality Circles work for General Practitioners - a realist approach 93%
- Surgery & COVID-19: A rapid scoping review of the impact of COVID-19 on surgical services during public health emergencies 93%
Similar papers in this journal
- A typology of physician input approaches to using AI chatbots for clinical decision-making: a mixed methods study 93%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 92%
- Can co-designed educational interventions help consumers think critically about asking ChatGPT health questions? Results from a randomised-controlled trial 92%
Similar papers in this journal
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 93%
- The NASSS (Non-Adoption, Abandonment, Scale-Up, Spread and Sustainability) framework use over time: A scoping review 93%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 92%
Similar papers in this journal
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 92%
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 92%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
Similar papers in this journal
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 93%
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 93%
- Artificial Intelligence (AI)-based Chatbots in Promoting Health Behavioral Changes: A Systematic Review 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.