Conversational multi-turn interaction does not ensure triage-disposition alignment in ChatGPT Health in real and synthetic patient encounters
Tyagi, A.; Taylor, B.; Jayaraman, P.; Jangda, M.; Jiang, J.; Kavula, S.; Gorin, M. A.; Hugo, H.; Gavin, N.; Freeman, R.; Charney, A. W.; Sakhuja, A.; Luo, Y.; Stump, L.; Ramaswamy, A.; Nadkarni, G. N.; Naved, B. A.
Show abstract
Individuals increasingly use conversational AI systems for symptom guidance. Whether multi-turn interactions improve clinical triage standard alignment remains uncertain. We conducted a retrospective, cross-sectional evaluation of 255 cases from three physician-reviewed sources: clinically authored vignettes (n=39), and real-world emergency department (N=76) and nurse line cases (n=140). CGPTH single-turn generated triage recommendations after only receiving an initial symptom description, reflecting potential typical use. Second, CGPTH multi-turn simulated nurse triage by asking subsequent questions before generating triage recommendations. Against nurse line standards, 52.9% single-turn and 55.7% multi-turn use prompting agreed exactly; clinician adjudicated disposition agreement was 54.1% and 48.2%. Discordant case recommendations represented lower acuity against nurse triage (natural: 70.8% under-triage, P < 0.0001; multi-turn: 69.0%, P < 0.0001). These findings suggest conversational interaction does not ensure safe triage-disposition alignment. Alongside aggregate agreement, clinical AI systems evaluations for symptom guidance should measure ordinal distance from standards, error direction, and additional dialogue conditions that affect recommended care standards.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 93%
- A Framework to Assess Clinical Safety and Hallucination Rates of LLMs for Medical Text Summarisation 92%
- Bridging the Literacy Gap for Surgical Consents: An AI-Human Expert Collaborative Approach 92%
Similar papers in this journal
- Observer: Creation of a Novel Multimodal Dataset for Outpatient Care Research 94%
- Clinical Utility of Automatable Prediction Models for Improving Palliative and End-Of-Life Care Outcomes: Towards Routine Decision Analysis Before Implementation 93%
- Real-Time Electronic Health Record Mortality Prediction During the COVID-19 Pandemic: A Prospective Cohort Study 93%
Similar papers in this journal
- Low adherence to existing model reporting guidelines by commonly used clinical prediction models 93%
- Score for Emergency Risk Prediction (SERP): An Interpretable Machine Learning AutoScore–Derived Triage Tool for Predicting Mortality after Emergency Admissions 92%
- Diagnostic Codes in AI prediction models and Label Leakage of Same-admission Clinical Outcomes 90%
Similar papers in this journal
- User Testing of a Diagnostic Decision Support System with Machine-assisted Chart Review to Facilitate Clinical Genomic Diagnosis 91%
- Natural Language Word-Embeddings as a glimpse into healthcare at the End Of Life 90%
- The performance of national COVID-19 ‘Symptom Checkers’: A comparative case simulation study 90%
Similar papers in this journal
- Real-world evaluation of AI-driven COVID-19 triage for emergency admissions: External validation & operational assessment of lab-free and high-throughput screening solutions 92%
- Remote Covid Assessment in Primary Care (RECAP) risk prediction tool: derivation and real-world validation studies 91%
- Automated and partially-automated contact tracing: a rapid systematic review to inform the control of COVID-19 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.