Accuracy and Consistency of Online Chat-based Artificial Intelligence Platforms in Answering Patients Questions on Heart Failure
Kozaily, E.; Geagea, M.; Akdogan, E. R.; Atkins, J.; Elshazly, M. B.; Guglin, M.; Tedford, R. J.; Wehbe, R. M.
Show abstract
BackgroundHeart failure (HF) is a prevalent condition associated with significant morbidity. Patients may have questions that they feel embarrassed to ask or will face delays awaiting responses from their healthcare providers which may impact their health behavior. We aimed to investigate the potential of chat-based artificial intelligence (AI) platforms in complementing the delivery of patient-centered care. MethodsUsing online patient forums and physician experience, we created 30 questions related to diagnosis, management and prognosis of HF. The questions were posed to two artificial intelligence (AI) chatbots (OpenAIs ChatGPT-3.5 and Googles Bard). Each set of answers was evaluated by two HF experts, independently and blinded to each other, for accuracy (adequacy of content) and consistency of content. ResultsChatGPT provided mostly appropriate answers (27/30, 90%) and showed a high degree of consistency (93%). Bard provided a similar content in its answers and thus was evaluated only for adequacy (23/30, 77%). The two HF experts grades were concordant in 83% and 67% of the questions for ChatGPT and Bard, respectively. Both platforms suffered from issues related to "hallucination" of facts and/or difficulty with more contemporary recommendations. ConclusionAI based chatbots may have potential in improving HF education and empowering patients, but their limitations should be considered and addressed in future research.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Cardiology Knowledge Assessment of Retrieval-Augmented Open versus Proprietary Large Language Models 93%
- Use of a Continuous Single Lead Electrocardiogram Analytic to Predict Patient Deterioration Requiring Rapid Response Team Activation 92%
- QRS detection in single-lead, telehealth electrocardiogram signals: benchmarking open-source algorithms 91%
Similar papers in this journal
- Integrating remote monitoring into heart failure patients’ care regimen: A pilot study 94%
- Predicting 30-Day and 1-Year Mortality in Heart Failure with Preserved Ejection Fraction (HFpEF) 94%
- Perceived barriers and enablers influencing physical activity in heart failure: a qualitative one-to-one interview study 94%
Similar papers in this journal
- Remote patient monitoring and digital therapeutics in heart failure: lessons from the Continuum pilot study 96%
- Improving emergency department patient-doctor conversation through an artificial intelligence symptom taking tool: an action-oriented design pilot study 90%
- The Validity of the Parsley Symptom Index: an e-PROM designed for Telehealth 90%
Similar papers in this journal
- A digital self-care intervention for Ugandan patients with heart failure and their clinicians: User-centred design and usability study 94%
- Validating a Clinical Decision Support System for Palliative Care using healthcare professionals’ insights 92%
- How suitable are clinical vignettes for the evaluation of symptom checker apps? A test theoretical perspective 89%
Similar papers in this journal
- Automated Diagnostic Reports from Images of Electrocardiograms at the Point-of-Care 92%
- Development and Validation of a parsimonious AI-Based Risk Score for Mortality in Heart Failure: A UK cohort study 90%
- Automated Echocardiographic Detection of Mitral Valve Prolapse and Mitral Regurgitation with Video-based Artificial Intelligence Algorithms 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.