Back

Evaluating ChatGPT's Performance in Responding to Questions About Endoscopic Procedures for Patients

Ali, H.; Patel, P.; Obaitan, I.; Mohan, B. P.; Sohail, A. H.; Smith-Martinez, L.; Lambert, K.; Gangwani, M. K.; Easler, J. J.; Adler, D. G.

2023-06-05 gastroenterology
10.1101/2023.05.31.23290800 medRxiv
Show abstract

Background and aimsWe aimed to assess the accuracy, completeness, and consistency of ChatGPTs responses to frequently asked questions concerning the management and care of patients receiving endoscopic procedures and to compare its performance to Generative Pre-trained Transformer 4 (GPT-4) in providing emotional support. MethodsFrequently asked questions (N = 117) about esophagogastroduodenoscopy (EGD), colonoscopy, endoscopic ultrasound (EUS), and endoscopic retrograde cholangiopancreatography (ERCP) were collected from professional societies, institutions, and social media. ChatGPTs responses were generated and graded by board-certified gastroenterologists and advanced endoscopists. Emotional support questions were assessed by a psychiatrist. ResultsChatGPT demonstrated high accuracy in answering questions about EGD (94.8% comprehensive or correct but insufficient), colonoscopy (100% comprehensive or correct but insufficient), ERCP (91% comprehensive or correct but insufficient), and EUS (87% comprehensive or correct but insufficient). No answers were deemed entirely incorrect (0%). Reproducibility was significant across all categories. ChatGPTs emotional support performance was inferior to the newer GPT-4 model. ConclusionChatGPT provides accurate and consistent responses to patient questions about common endoscopic procedures and demonstrates potential as a supplementary information resource for patients and healthcare providers.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.