Back

ChatIBD: design, safeguards, and early international use of a guideline-grounded generative AI tool for inflammatory bowel disease (IBD) professionals

Chuah, C. S.; Gros, B.; Plevris, N.

2026-05-07 gastroenterology
10.64898/2026.05.06.26352526 medRxiv
Show abstract

ObjectivesTo describe the design, operational safeguards, and early use of ChatIBD, a specialty-specific generative AI platform for inflammatory bowel disease (IBD), during its first 6 months of live deployment. MethodsChatIBD is an online question-answering platform that uses retrieval-augmented generation over a curated corpus of IBD guidelines. Queries undergo hybrid semantic and keyword retrieval with query expansion and reranking, and the model is instructed to answer only from retrieved material and return linked citations. Safeguards include fixed medication dosing information from European Medicines Agency (EMA), user feedback capture, and clinician review of flagged outputs. We performed a descriptive service evaluation of aggregated, de-identified platform metrics collected between 1 October 2025 and 1 April 2026. ResultsDuring the study period, ChatIBD registered 913 users and processed 7,222 messages across 3,855 conversations. Activity was recorded across 69 countries and 28 languages, with the highest message volumes from the United Kingdom (27.1%) and Spain (12.3%). Median daily message volume was 35.5 (IQR 20 to 52), and 85.1% of messages were submitted on weekdays. Medication-related queries accounted for the largest use domain, while guideline synthesis was the most frequent inferred intent. Sixteen explicit feedback events were recorded, including one negative rating that triggered clinician review and system changes. ConclusionsChatIBD showed early international uptake and repeat use as a specialty-specific, retrieval-grounded generative AI tool for IBD professionals. These findings support the feasibility of deploying a guideline-grounded clinical AI service with practical safeguards, but do not establish response accuracy, safety, or clinical effectiveness. Formal validation is in progress. What is already known on this topicGeneral-purpose large language models are increasingly being used informally by clinicians, but concerns remain about hallucinated content, unverifiable recommendations, and poor traceability to specialty-specific sources. What this study addsThis study describes the early deployment of ChatIBD, a specialty-specific retrieval-grounded generative AI tool for IBD professionals, and the safeguards used in its live operation. How this study might affect research, practice or policyEarly evaluations of live specialist AI tools may help guide governance, implementation, and validation. Uptake alone is not evidence of effectiveness, but it can help shape priorities for subsequent studies.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.