Back

Accuracy and Consistency of Online Chat-based Artificial Intelligence Platforms in Answering Patients Questions on Heart Failure

Kozaily, E.; Geagea, M.; Akdogan, E. R.; Atkins, J.; Elshazly, M. B.; Guglin, M.; Tedford, R. J.; Wehbe, R. M.

2023-09-13 cardiovascular medicine
10.1101/2023.09.12.23295452 medRxiv
Show abstract

BackgroundHeart failure (HF) is a prevalent condition associated with significant morbidity. Patients may have questions that they feel embarrassed to ask or will face delays awaiting responses from their healthcare providers which may impact their health behavior. We aimed to investigate the potential of chat-based artificial intelligence (AI) platforms in complementing the delivery of patient-centered care. MethodsUsing online patient forums and physician experience, we created 30 questions related to diagnosis, management and prognosis of HF. The questions were posed to two artificial intelligence (AI) chatbots (OpenAIs ChatGPT-3.5 and Googles Bard). Each set of answers was evaluated by two HF experts, independently and blinded to each other, for accuracy (adequacy of content) and consistency of content. ResultsChatGPT provided mostly appropriate answers (27/30, 90%) and showed a high degree of consistency (93%). Bard provided a similar content in its answers and thus was evaluated only for adequacy (23/30, 77%). The two HF experts grades were concordant in 83% and 67% of the questions for ChatGPT and Bard, respectively. Both platforms suffered from issues related to "hallucination" of facts and/or difficulty with more contemporary recommendations. ConclusionAI based chatbots may have potential in improving HF education and empowering patients, but their limitations should be considered and addressed in future research.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.