OpenEvidence errs on the safe side in a structured test of triage recommendations
Jia, E.; Omar, M.; Barash, Y.; Brook, O. R.; Ahmed, M.; Kruskal, J. B.; Gorenshtein, A.; Klang, E.
Show abstract
Ramaswamy et al. recently reported in Nature Medicine that ChatGPT Health, a consumer-facing health AI tool, undertriaged 51.6% of true emergencies. It was also susceptible to social anchoring in a structured stress test of triage recommendations. We applied the same vignette-based benchmark to OpenEvidence, a widely used physician-facing AI platform for clinical decision support. The benchmark included 960 prompts across 21 clinical domains (Supplementary Table S3). OpenEvidence undertriaged 12.5% of emergencies, a four-fold reduction relative to ChatGPT Health. It also showed no anchoring effect. Its errors skewed in a safer direction, including 68.0% overtriage of Home presentations. In 65 of 960 responses (6.8%), it declined to assign a triage level. These refusals occurred only in symptom-only prompts and never in urgent or emergency cases. Performance improved when objective clinical data were provided. Under the same benchmark, a widely used physician-facing system showed a different safety profile from a consumer-facing one. This suggests that who a health AI is built for can shape how it fails.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Finding Long-COVID: Temporal Topic Modeling of Electronic Health Records from the N3C and RECOVER Programs 93%
- Development and assessment of a machine learning tool for predicting emergency admission in Scotland 92%
- Machine Learning for Real-Time Aggregated Prediction of Hospital Admission for Emergency Patients 91%
Similar papers in this journal
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 94%
- Understanding COVID-19 trajectories from a nationwide linked electronic health record cohort of 56 million people: phenotypes, severity, waves & vaccination 93%
- Predicting hospital-onset COVID-19 infections using dynamic networks of patient contacts: an observational study 92%
Similar papers in this journal
- Evaluating and Mitigating Limitations of Large Language Models in Clinical Decision Making 91%
- Attributes and predictors of Long-COVID: analysis of COVID cases and their symptoms collected by the Covid Symptoms Study App 90%
- Zero-shot drug repurposing with geometric deep learning and clinician centered design 89%
Similar papers in this journal
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 92%
- Collaborative Large Language Models for Automated Data Extraction in Living Systematic Reviews 91%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 91%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 90%
- Effectiveness of the Single-Dose Ad26.COV2.S COVID Vaccine 89%
- Clinical Outcomes, Costs, and Cost-effectiveness of Strategies for People Experiencing Sheltered Homelessness During the COVID-19 Pandemic 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.