Back

Comparison of ChatGPT vs. Bard to Anesthesia-related Queries

Patnaik, S. S.; Hoffmann, U.

2023-06-30 anesthesia
10.1101/2023.06.29.23292057 medRxiv
Show abstract

We investigated the ability of large language models (LLMs) to answer anesthesia related queries prior to surgery from a patients point of view. In the study, we introduced textual data evaluation metrics, investigated "hallucinations" phenomenon, and evaluated feasibility of using LLMs at the patient-clinician interface. ChatGPT was found to be lengthier, intellectual, and effective in its response as compared to Bard. Upon clinical evaluation, no "hallucination" errors were reported from ChatGPT, whereas we observed a 30.3% error in response from Bard. ChatGPT responses were difficult to read (college level difficulty) while Bard responses were more conversational and about 8th grade level from readability calculations. Linguistic quality of ChatGPT was found to be 19.7% greater for Bard (66.16 {+/-} 13.42 vs. 55.27 {+/-} 11.76; p=0.0037) and was independent of response length. Computational sentiment analysis revelated that polarity scores of on a Bard was significantly greater than ChatGPT (mean 0.16 vs. 0.11 on scale of -1 (negative) to 1 (positive); p=0.0323) and can be classified as "positive"; whereas subjectivity scores were similar across LLMs (mean 0.54 vs 0.50 on a scale of 0 (objective) to 1 (subjective), p=0.3030). Even though the majority of the LLM responses were appropriate, at this stage these chatbots should be considered as a versatile clinical resource to assist communication between clinicians and patients, and not a replacement of essential pre-anesthesia consultation. Further efforts are needed to incorporate health literacy that will improve patient-clinical communications and ultimately, post-operative patient outcomes.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.