Evaluating self-triage accuracy of laypeople, symptom-assessment apps, and large language models: A framework for case vignette development using a representative design approach (RepVig)
Kopka, M.; Napierala, H.; Privoznik, M.; Sapunova, D.; Zhang, S.; Feufel, M.
Show abstract
Most studies evaluating symptom-assessment applications (SAAs) rely on a common set of case vignettes that are authored by clinicians and devoid of context, which may be representative of clinical settings but not of situations where patients use SAAs. Assuming the use case of self-triage, we used representative design principles to sample case vignettes from online platforms where patients describe their symptoms to obtain professional advice and compared triage performance of laypeople, SAAs, and Large Language Models (LLMs) on representative versus standard vignettes. We found performance differences in all three groups depending on vignette type (OR = 1.27 to 3.41, p < .001 to .035) and changed rankings of best-performing SAAs and LLMs. Based on these results, we argue that our representative vignette sampling approach (that we call the RepVig Framework) should replace the practice of using a fixed vignette set as standard for SAA evaluation studies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- A proposed de-identification framework for a cohort of children presenting at a health facility in Uganda 94%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 93%
- Predictability and Stability Testing to Assess Clinical Decision Instrument Performance for Children After Blunt Torso Trauma 93%
Similar papers in this journal
- Measures of socioeconomic advantage are not independent predictors of support for healthcare AI: subgroup analysis of a national Australian survey 93%
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 92%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 92%
Similar papers in this journal
- What is the suitability of clinical vignettes in benchmarking the performance of online symptom checkers? An audit study 95%
- How and why do Quality Circles work for General Practitioners - a realist approach 94%
- Trans-sectoral patient pathways in urgent and emergency care: a Study Protocol for a prospective mixed-methods study in Germany (TRANSPARENT Study) 93%
Similar papers in this journal
- Using explainable machine learning to identify patients at risk of reattendance at discharge from emergency departments 93%
- The impact of atypical intrahospital transfers on patient outcomes: a mixed methods study 92%
- Scalable Incident Detection via Natural Language Processing and Probabilistic Language Models 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.