Back

Examining Batch Effect in Histopathology as a Distributionally Robust Optimization Problem

Hari, S. N.; Nyman, J.; Mehta, N.; Jiang, B.; Rosenthal, J.; Sengupta, E.; Dietlein, F.; Umeton, R.; Van Allen, E. M.

2021-09-15 cancer biology
10.1101/2021.09.14.460365 bioRxiv
Show abstract

Computer vision (CV) approaches applied to digital pathology have informed biological discovery and development of tools to help inform clinical decision-making. However, batch effects in the images have the potential to introduce spurious confounders and represent a major challenge to effective analysis and interpretation of these data. Standard methods to circumvent learning such confounders include (i) application of image augmentation techniques and (ii) examination of the learning process by evaluating through external validation (e.g., unseen data coming from a comparable dataset collected at another hospital). Here, we show that the source site of a histopathology slide can be learned from the image using CV algorithms in spite of image augmentation, and we explore these source site predictions using interpretability tools. A CV model trained using Empirical Risk Minimization (ERM) risks learning this source-site signal as a spurious correlate in the weak-label regime, which we abate by using a training method with abstention. We find that a patch based classifier trained using abstention outperformed a model trained using ERM by 9.9, 10 and 19.4% F1 in the binary classification tasks of identifying tumor versus normal tissue in lung adenocarcinoma, Gleason score in prostate adenocarcinoma, and tumor tissue grade in clear cell renal cell carcinoma, respectively, at the expense of up to 80% coverage (defined as the percent of tiles not abstained on by the model). Further, by examining the areas abstained by the model, we find that the model trained using abstention is more robust to heterogeneity, artifacts and spurious correlates in the tissue. Thus, a method trained with abstention may offer novel insights into relevant areas of the tissue contributing to a particular phenotype. Together, we suggest using data augmentation methods that help mitigate a digital pathology models reliance on potentially spurious visual features, as well as selecting models that can identify features truly relevant for translational discovery and clinical decision support.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
PLOS Computational Biology
1863 papers in training set
Top 1%
18.9%
2
PLOS ONE
5266 papers in training set
Top 14%
13.2%
3
Journal of Medical Imaging
11 papers in training set
Top 0.1%
11.3%
4
Scientific Reports
3612 papers in training set
Top 10%
6.9%
50% of probability mass above
5
npj Digital Medicine
118 papers in training set
Top 1%
4.9%
6
Patterns
78 papers in training set
Top 0.4%
3.6%
7
JCO Clinical Cancer Informatics
22 papers in training set
Top 0.2%
3.6%
8
Modern Pathology
22 papers in training set
Top 0.1%
2.7%
9
PLOS Digital Health
106 papers in training set
Top 2%
1.9%
10
Computers in Biology and Medicine
128 papers in training set
Top 2%
1.9%
11
Computer Methods and Programs in Biomedicine
28 papers in training set
Top 0.4%
1.8%
12
Cancers
213 papers in training set
Top 3%
1.8%
13
Frontiers in Bioinformatics
49 papers in training set
Top 0.5%
1.5%
14
npj Precision Oncology
53 papers in training set
Top 0.9%
1.5%
15
Biology Methods and Protocols
61 papers in training set
Top 0.9%
1.5%
16
Medical Physics
14 papers in training set
Top 0.4%
1.2%
17
Nature Communications
5641 papers in training set
Top 50%
1.2%
18
GigaScience
212 papers in training set
Top 4%
1.1%
19
IEEE Transactions on Computational Biology and Bioinformatics
20 papers in training set
Top 0.5%
1.1%
20
Cancer Research Communications
51 papers in training set
Top 1%
1.0%
21
Medical Image Analysis
35 papers in training set
Top 0.7%
0.9%
22
Communications Medicine
113 papers in training set
Top 5%
0.9%
23
Journal of Translational Medicine
57 papers in training set
Top 2%
0.9%
24
npj Imaging
12 papers in training set
Top 0.2%
0.9%
25
JAMIA Open
42 papers in training set
Top 1%
0.9%
26
The American Journal of Pathology
32 papers in training set
Top 0.8%
0.6%
27
Nature Machine Intelligence
70 papers in training set
Top 3%
0.6%
28
Journal of Pathology Informatics
15 papers in training set
Top 0.3%
0.6%
29
Bioinformatics Advances
203 papers in training set
Top 5%
0.6%
30
BMC Medical Informatics and Decision Making
43 papers in training set
Top 2%
0.6%