Back

Impact of Whole Slide Image Blurriness on the Robustness of Artificial Intelligence in Real World Setting: Retrospective Observational Study

Kim, H. H.; Ko, Y. S.; Kim, K.

2025-03-06 pathology
10.1101/2025.03.05.25323268 medRxiv
Show abstract

ContextIn digital pathology, blurriness in whole slide images (WSI) is a common issue, with severe blurriness widely acknowledge as a critical factor that can degrade the performance of artificial intelligence (AI) models. However, the effects of the typical levels of blurriness observed in real-world pathological images on the robustness of AI predictions remains unclear and unexplored. ObjectiveTo evaluate the impact of WSI blurring on the robustness of AI prediction in real-world setting. DesignA retrospective study was conducted using 8,000 WSIs and corresponding AI predictions from four AI models trained on data from two scanners and two organs. WSIs were categorized into concordant and discordant groups based on AI-prediction accuracy. Analyses included: 1) comparing blur metrics between groups, 2) determining the odds ratio between the proportions of blurry patch in WSIs and prediction concordance, and 3) assessing model performance across varying blur intensities. ResultsFor each organ-scanner pair, the average wavelet score and Laplacian variance for WSIs between the two groups did not show a statistically significant difference model (p > 0.05 for both metrics), except for one, and their effect sizes were small (Cohens D < 0.2 for both metrics). Additionally, no statistically significant association was observed between AI prediction concordance and the proportion of blurry images in WSIs (confidence intervals included 1, respectively). Model performance remained robust even at high blur level (radius=1) at which patch image had Laplacian variance of 162.88 and a wavelet score of 1880.07, corresponding to the top 1.22% and 2.16% of blurriness respective, in our dataset. ConclusionsThe findings empirically suggest that the typical levels of WSI blurriness encountered in real-world settings may not significantly compromise the robustness of AI predictions.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Journal of Pathology Informatics
15 papers in training set
Top 0.1%
30.7%
2
Scientific Reports
3612 papers in training set
Top 8%
7.8%
3
npj Digital Medicine
118 papers in training set
Top 0.8%
7.2%
4
Modern Pathology
22 papers in training set
Top 0.1%
6.2%
50% of probability mass above
5
Diagnostics
50 papers in training set
Top 0.3%
5.4%
6
The Lancet Digital Health
25 papers in training set
Top 0.1%
5.4%
7
Journal of Medical Imaging
11 papers in training set
Top 0.1%
5.1%
8
PLOS ONE
5266 papers in training set
Top 36%
3.4%
9
The American Journal of Pathology
32 papers in training set
Top 0.2%
3.1%
10
eBioMedicine
183 papers in training set
Top 2%
2.4%
11
npj Precision Oncology
53 papers in training set
Top 0.8%
1.9%
12
JAMA Network Open
130 papers in training set
Top 2%
1.7%
13
Communications Medicine
113 papers in training set
Top 4%
1.1%
14
Frontiers in Medicine
120 papers in training set
Top 3%
1.1%
15
Laboratory Investigation
13 papers in training set
Top 0.2%
1.0%
16
Medical Image Analysis
35 papers in training set
Top 0.6%
1.0%
17
Computers in Biology and Medicine
128 papers in training set
Top 4%
1.0%
18
Biology Methods and Protocols
61 papers in training set
Top 2%
0.9%
19
Nature Communications
5641 papers in training set
Top 57%
0.8%
20
IEEE Access
35 papers in training set
Top 1%
0.8%
21
Cancer Research Communications
51 papers in training set
Top 2%
0.8%
22
The Journal of Molecular Diagnostics
39 papers in training set
Top 0.7%
0.6%
23
Advanced Intelligent Systems
11 papers in training set
Top 0.4%
0.6%
24
Breast Cancer Research
36 papers in training set
Top 0.7%
0.6%
25
BMC Cancer
67 papers in training set
Top 2%
0.6%
26
Scientific Data
209 papers in training set
Top 3%
0.6%