Screening of normal endoscopic large bowel biopsies with artificial intelligence: a retrospective study
Graham, S.; Minhas, F.; Bilal, M.; Ali, M.; Tsang, Y. W.; Eastwood, M.; Wahab, N.; Jahanifar, M.; Hero, E.; Dodd, K.; Sahota, H.; Wu, S.; Lu, W.; Azam, A.; Benes, K.; Nimir, M.; Hewitt, K.; Bhalerao, A.; Robinson, A.; Eldaly, H.; E Ahmed Raza, S.; Gopalakrishnan, K.; Snead, D.; Rajpoot, N.
Show abstract
ObjectivesDevelop an interpretable AI algorithm to rule out normal large bowel endoscopic biopsies saving pathologist resources. DesignRetrospective study. SettingOne UK NHS site was used for model training and internal validation. External validation conducted on data from two other NHS sites and one site in Portugal. Participants6,591 whole-slides images of endoscopic large bowel biopsies from 3,291 patients (54% Female, 46% Male). Main outcome measuresArea under the receiver operating characteristic and precision recall curves (AUC-ROC and AUC-PR), measuring agreement between consensus pathologist diagnosis and AI generated classification of normal versus abnormal biopsies. ResultsA graph neural network was developed incorporating pathologist domain knowledge to classify the biopsies as normal or abnormal using clinically driven interpretable features. Model training and internal validation were performed on 5,054 whole slide images of 2,080 patients from a single NHS site resulting in an AUC-ROC of 0.98 (SD=0.004) and AUC-PR of 0.98 (SD=0.003). The predictive performance of the model was consistent in testing over 1,537 whole slide images of 1,211 patients from three independent external datasets with mean AUC-ROC = 0.97 (SD=0.007) and AUC-PR = 0.97 (SD=0.005). Our analysis shows that at a high sensitivity threshold of 99%, the proposed model can, on average, reduce the number of normal slides to be reviewed by a pathologist by 55%. A key advantage of IGUANA is its ability to provide an explainable output highlighting potential abnormalities in a whole slide image as a heatmap overlay in addition to numerical values associating model prediction with various histological features. Example results with can be viewed online at https://iguana.dcs.warwick.ac.uk/. ConclusionsAn interpretable AI model was developed to screen abnormal cases for review by pathologists. The model achieved consistently high predictive accuracy on independent cohorts showing its potential in optimising increasingly scarce pathologist resources and for achieving faster time to diagnosis. Explainable predictions of IGUANA can guide pathologists in their diagnostic decision making and help boost their confidence in the algorithm, paving the way for future clinical adoption. What is already known on this topicO_LIIncreasing screening rates for early detection of colon cancer are placing significant pressure on already understaffed and overloaded histopathology resources worldwide and especially in the United Kingdom1. C_LIO_LIApproximately a third of endoscopic colon biopsies are reported as normal and therefore require minimal intervention, yet the biopsy results can take up to 2-3 weeks2. C_LIO_LIAI models hold great promise for reducing the burden of diagnostics for cancer screening but require incorporation of pathologist domain knowledge and explainability. C_LI What this study addsO_LIThis study presents the first AI algorithm for rule out of normal from abnormal large bowel endoscopic biopsies with high accuracy across different patient populations. C_LIO_LIFor colon biopsies predicted as abnormal, the model can highlight diagnostically important biopsy regions and provide a list of clinically meaningful features of those regions such as glandular architecture, inflammatory cell density and spatial relationships between inflammatory cells, glandular structures and the epithelium. C_LIO_LIThe proposed tool can both screen out normal biopsies and act as a decision support tool for abnormal biopsies, therefore offering a significant reduction in the pathologist workload and faster turnaround times. C_LI
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 97%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 95%
- CARDBiomedBench: A Benchmark for Evaluating Large Language Model Performance in Biomedical Research 90%
Similar papers in this journal
- A tool for federated training of segmentation models on whole slide images. 96%
- Using an Anomaly Detection Approach for the Segmentation of Colorectal Cancer Tumors in Whole Slide Images 95%
- Development of an Interactive Web Dashboard to Facilitate the Reexamination of Pathology Reports for Instances of Underbilling of CPT Codes 94%
Similar papers in this journal
Similar papers in this journal
- A human-in-the-loop explanation framework for morphologically transparent AI predictions from whole-slide images 94%
- A Deep Learning Based Smartphone Application for Early Detection of Nasopharyngeal Carcinoma Using Endoscopic Images 92%
- High-Sensitivity Pan-Cancer AI Assessment of Lymph Node Metastasis via Uncertainty Quantification 92%
Similar papers in this journal
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 97%
- Interpretable multimodal deep learning for real-time pan-tissue pan-disease pathology search on social media 94%
- Clinical-Grade Validation of an Autofluorescence Virtual Staining System with Human Experts and a Deep Learning System for Prostate Cancer 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.