Comparison of ChatGPT vs. Bard to Anesthesia-related Queries
Patnaik, S. S.; Hoffmann, U.
Show abstract
We investigated the ability of large language models (LLMs) to answer anesthesia related queries prior to surgery from a patients point of view. In the study, we introduced textual data evaluation metrics, investigated "hallucinations" phenomenon, and evaluated feasibility of using LLMs at the patient-clinician interface. ChatGPT was found to be lengthier, intellectual, and effective in its response as compared to Bard. Upon clinical evaluation, no "hallucination" errors were reported from ChatGPT, whereas we observed a 30.3% error in response from Bard. ChatGPT responses were difficult to read (college level difficulty) while Bard responses were more conversational and about 8th grade level from readability calculations. Linguistic quality of ChatGPT was found to be 19.7% greater for Bard (66.16 {+/-} 13.42 vs. 55.27 {+/-} 11.76; p=0.0037) and was independent of response length. Computational sentiment analysis revelated that polarity scores of on a Bard was significantly greater than ChatGPT (mean 0.16 vs. 0.11 on scale of -1 (negative) to 1 (positive); p=0.0323) and can be classified as "positive"; whereas subjectivity scores were similar across LLMs (mean 0.54 vs 0.50 on a scale of 0 (objective) to 1 (subjective), p=0.3030). Even though the majority of the LLM responses were appropriate, at this stage these chatbots should be considered as a versatile clinical resource to assist communication between clinicians and patients, and not a replacement of essential pre-anesthesia consultation. Further efforts are needed to incorporate health literacy that will improve patient-clinical communications and ultimately, post-operative patient outcomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An interactive retrieval system for clinical trial studies with context-dependent protocol elements 93%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 93%
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 93%
Similar papers in this journal
Similar papers in this journal
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 93%
- Design and implementation of a system for automated monitoring of adherence to evidenced-based clinical guideline recommendations 93%
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 93%
Similar papers in this journal
- Comparative Analysis of Multimodal Large Language Models GPT-4o and o1 vs Clinicians in Clinical Case Challenge Questions 93%
- Open Science Practices Among Authors Published in Complementary, Alternative, and Integrative Medicine Journals: An International, Cross-Sectional Survey 91%
- The PBL teaching method in Neurology Education in the Traditional Chinese Medicine undergraduate students: An Observational Study 91%
Similar papers in this journal
- Performance of o1 pro and GPT-4 in self-assessment questions for nephrology board renewal 91%
- Emerging Applications of NLP and Large Language Models in Gastroenterology and Hepatology: A Systematic Review 91%
- Streamlining Intersectoral Provision of Real-World Health Data: A Service Platform for Improved Clinical Research and Patient Care 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.