Back

Design and Implementation of an End-to-End AI-Driven Colonoscopy Recall Workflow at Scale

Mohapatra, A.; Porth, R.; Wong, S.; Hardy, H.; Piatkowski, G.; Shang, J.; Amat, M.; Flier, S.; Salsman, A.; Fitzgerald, T.; Shammout, A.; Rubins, D.; Miller, A.; Jegadeesan, V.; Ravi, A.; Feuerstein, J.

2025-07-14 gastroenterology
10.1101/2025.07.11.25331400 medRxiv
Show abstract

Article DescriptionWe present a real-world deployment of a large language model-powered colonoscopy recall pipeline that structured over 100,000 patient records during an EHR transition. This end-to-end AI system demonstrated high fidelity, scalability, and a projected prevention of up to 6092 colorectal cancer cases and cost savings between 400 - 670 million dollars. The rise of structured data elements in Electronic Health Records (EHRs) is a key enabler of improving care quality. However, the transition towards routine use of these fields paradoxically heightens patient safety risks due to increased variability in documentation and the use of "placeholder" values pending manual review. For large clinical initiatives such as colon cancer screening and surveillance, misinterpretation of recorded clinical data can be particularly problematic, disrupting risk-adapted recall guidance and potentially exacerbating care gaps. This case study details the development and deployment of a Large Language Model (LLM)-driven workflow to extract and transfer unstructured colonoscopy recall recommendations as part of a larger EHR migration. Utilizing GPT-4 Turbo for the core inference step of a fully integrated pipeline--spanning custom SQL queries, Optical Character Recognition (OCR) of historical PDFs, LLM-based inference, and anomaly detection--we successfully structured and migrated population-wide colonoscopy recall data corresponding to over 100,00 patients and 10 years of clinical care. The pipeline demonstrated high accuracy (Macro F1=1.0 against clinician review), scalability, and cost efficiency. We estimate that use of this workflow--relative to the alternative of a default 10-year reminder from last colonoscopy--may prevent over 6,000 new colorectal cancer cases (a projected cost savings of $400-670 million). Key lessons from implementation include the importance of stakeholder alignment, the necessity of robust quality control at scale, and the technical challenges of expanding optimized LLM inference to a fully-fledged end-to-end clinical workflow.

Published in JAMIA Open (predicted rank #1) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.