Impact of Whole Slide Image Blurriness on the Robustness of Artificial Intelligence in Real World Setting: Retrospective Observational Study
Kim, H. H.; Ko, Y. S.; Kim, K.
Show abstract
ContextIn digital pathology, blurriness in whole slide images (WSI) is a common issue, with severe blurriness widely acknowledge as a critical factor that can degrade the performance of artificial intelligence (AI) models. However, the effects of the typical levels of blurriness observed in real-world pathological images on the robustness of AI predictions remains unclear and unexplored. ObjectiveTo evaluate the impact of WSI blurring on the robustness of AI prediction in real-world setting. DesignA retrospective study was conducted using 8,000 WSIs and corresponding AI predictions from four AI models trained on data from two scanners and two organs. WSIs were categorized into concordant and discordant groups based on AI-prediction accuracy. Analyses included: 1) comparing blur metrics between groups, 2) determining the odds ratio between the proportions of blurry patch in WSIs and prediction concordance, and 3) assessing model performance across varying blur intensities. ResultsFor each organ-scanner pair, the average wavelet score and Laplacian variance for WSIs between the two groups did not show a statistically significant difference model (p > 0.05 for both metrics), except for one, and their effect sizes were small (Cohens D < 0.2 for both metrics). Additionally, no statistically significant association was observed between AI prediction concordance and the proportion of blurry images in WSIs (confidence intervals included 1, respectively). Model performance remained robust even at high blur level (radius=1) at which patch image had Laplacian variance of 162.88 and a wavelet score of 1880.07, corresponding to the top 1.22% and 2.16% of blurriness respective, in our dataset. ConclusionsThe findings empirically suggest that the typical levels of WSI blurriness encountered in real-world settings may not significantly compromise the robustness of AI predictions.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Using an Anomaly Detection Approach for the Segmentation of Colorectal Cancer Tumors in Whole Slide Images 96%
- Independent assessment of a deep learning system for lymph node metastasis detection on the Augmented Reality Microscope 95%
- Bladder Cancer Prognosis Using Deep Neural Networks and Histopathology Images 94%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 96%
- Clinical-Grade Validation of an Autofluorescence Virtual Staining System with Human Experts and a Deep Learning System for Prostate Cancer 95%
- MIXTURE of human expertise and deep learning—Developing an explainable model for predicting pathological diagnosis and survival in patients with interstitial lung disease 95%
Similar papers in this journal
- Deep learning models for poorly differentiated colorectal adenocarcinoma classification in whole slide images using transfer learning 95%
- Auto-detection of motion artifacts on CT pulmonary angiograms with a physician-trained AI algorithm 93%
- Demarcation line determination for diagnosis of gastric cancer disease range using unsupervised machine learning in magnifying narrow-band imaging 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.