ChatGPT Is Still Not Good Enough at Giving Care-Seeking Advice, or Is It?
Kopka, M.; He, L.; Feufel, M. A.
Show abstract
Artificial Intelligence tools like ChatGPT are increasingly used by patients to support their care-seeking decisions, although the accuracy of newer models remains unclear. We evaluated 16 ChatGPT models using 45 validated vignettes, each prompted ten times (7,200 total assessments). Each model classified the vignettes as requiring emergency care, non-emergency care, or self-care. We evaluated accuracy against each cases gold standard solution, examined the variability across trials, and tested algorithms to aggregate multiple recommendations to improve accuracy. o1-mini achieved the highest accuracy (78%), but we could not observe an overall improvement with newer models - although reasoning models (e.g., o4-mini) improved their accuracy in identifying self-care cases. Selecting the lowest urgency level across multiple trials improved accuracy by 4 percentage points. Although newer models slightly outperform laypeople, their accuracy remains insufficient for standalone use. However, making use of output variability with aggregation algorithms can improve the performance of these models.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 95%
- Harnessing the Open Access Version of ChatGPT for Enhanced Clinical Opinions 94%
- Development and preliminary testing of Health Equity Across the AI Lifecycle (HEAAL): A framework for healthcare delivery organizations to mitigate the risk of AI solutions worsening health inequities 94%
Similar papers in this journal
- Connecting Artificial Intelligence and Primary Care Challenges: Findings from a Multi-Stakeholder Collaborative Consultation 93%
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 93%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 92%
Similar papers in this journal
- Understanding how the design and implementation of Online Consultations influence primary care outcomes: Systematic review of evidence with recommendations for designers, providers, and researchers 93%
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.