Back

Deep learning detects virus presence in cancer histology

Kather, J. N.; Schulte, J.; Grabsch, H. I.; Loeffler, C.; Muti, H. S.; Dolezal, J.; Srisuwananukorn, A.; Agrawal, N.; Kochanny, S.; von Stillfried, S.; Boor, P.; Yoshikawa, T.; Jaeger, D.; Trautwein, C.; Bankhead, P.; Cipriani, N. A.; Luedde, T.; Pearson, A. T.

2019-07-05 cancer biology
10.1101/690206 bioRxiv
Show abstract

Oncogenic viruses like human papilloma virus (HPV) or Epstein Barr virus (EBV) are a major cause of human cancer. Viral oncogenesis has a direct impact on treatment decisions because virus-associated tumors can demand a lower intensity of chemotherapy and radiation or can be more susceptible to immune check-point inhibition. However, molecular tests for HPV and EBV are not ubiquitously available.\n\nWe hypothesized that the histopathological features of virus-driven and non-virus driven cancers are sufficiently different to be detectable by artificial intelligence (AI) through deep learning-based analysis of images from routine hematoxylin and eosin (HE) stained slides. We show that deep transfer learning can predict presence of HPV in head and neck cancer with a patient-level 3-fold cross validated area-under-the-curve (AUC) of 0.89 [0.82; 0.94]. The same workflow was used for Epstein-Barr virus (EBV) driven gastric cancer achieving a cross-validated AUC of 0.80 [0.70; 0.92] and a similar performance in external validation sets. Reverse-engineering our deep neural networks, we show that the key morphological features can be made understandable to humans.\n\nThis workflow could enable a fast and low-cost method to identify virus-induced cancer in clinical trials or clinical routine. At the same time, our approach for feature visualization allows pathologists to look into the black box of deep learning, enabling them to check the plausibility of computer-based image classification.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.