Evaluating ChatGPT's Performance in Responding to Questions About Endoscopic Procedures for Patients
Ali, H.; Patel, P.; Obaitan, I.; Mohan, B. P.; Sohail, A. H.; Smith-Martinez, L.; Lambert, K.; Gangwani, M. K.; Easler, J. J.; Adler, D. G.
Show abstract
Background and aimsWe aimed to assess the accuracy, completeness, and consistency of ChatGPTs responses to frequently asked questions concerning the management and care of patients receiving endoscopic procedures and to compare its performance to Generative Pre-trained Transformer 4 (GPT-4) in providing emotional support. MethodsFrequently asked questions (N = 117) about esophagogastroduodenoscopy (EGD), colonoscopy, endoscopic ultrasound (EUS), and endoscopic retrograde cholangiopancreatography (ERCP) were collected from professional societies, institutions, and social media. ChatGPTs responses were generated and graded by board-certified gastroenterologists and advanced endoscopists. Emotional support questions were assessed by a psychiatrist. ResultsChatGPT demonstrated high accuracy in answering questions about EGD (94.8% comprehensive or correct but insufficient), colonoscopy (100% comprehensive or correct but insufficient), ERCP (91% comprehensive or correct but insufficient), and EUS (87% comprehensive or correct but insufficient). No answers were deemed entirely incorrect (0%). Reproducibility was significant across all categories. ChatGPTs emotional support performance was inferior to the newer GPT-4 model. ConclusionChatGPT provides accurate and consistent responses to patient questions about common endoscopic procedures and demonstrates potential as a supplementary information resource for patients and healthcare providers.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using a Multilingual AI Care Agent to Reduce Disparities in Colorectal Cancer Screening: Higher FIT Test Adoption Among Spanish-Speaking Patients 93%
- Assessing ChatGPT’s Mastery of Bloom’s Taxonomy using psychosomatic medicine exam questions 93%
- Improving Patient Engagement in Phase 2 Clinical Trials with a Trial-specific Patient Decision Aid (tPDA): A Development and Usability Study 92%
Similar papers in this journal
- Prohibiting Babel - A call for professional remote interpreting services in pre-operation anaesthesia information 93%
- Protocol For Human Evaluation of Artificial Intelligence Chatbots in Clinical Consultations 92%
- Triaging and Referring In Adjacent General and Emergency Departments (the TRIAGE trial): a cluster randomised controlled trial 92%
Similar papers in this journal
- Development and Validation of the Alimetry(R) Gut-Brain Wellbeing Survey: A novel patient-reported mental health scale for patients with chronic gastroduodenal symptoms 93%
- A cross-sectional survey of material deprivation and suicide-related ideation among Vietnamese technical interns in Japan in 2021 87%
- The efficacy of a virtual reality exposure therapy treatment for fear of flying: A retrospective study 87%
Similar papers in this journal
- Effect of Virtually Led Value-Based Preoperative Assessment on Safety, Efficiency, and Patient and Professional Satisfaction 93%
- Artificial Intelligence in laryngeal endoscopy: Systematic Review and Meta-Analysis 93%
- Aerodigestoscopy (ADS): A retrospective examination of the feasibility, safety, and comfort of a new procedure for the evaluation of physiological disorders of the aerodigestive tract 92%
Similar papers in this journal
- Non-endoscopic screening for Barrett’s esophagus and Esophageal Adenocarcinoma in at risk Veterans 92%
- Colorectal cancer screening based on predicted risk: a pilot randomized controlled trial 92%
- Prune intake ameliorates chronic constipation symptoms and causes little discomfort from diarrhea and loose stools: A randomized placebo-controlled trial 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.