Back

Predicting the HER2 status in esophageal cancer from tissue microarrays using convolutional neural networks

Pisula, J. I.; Datta, R. R.; Boerner-Valdez, L.; Avemarg, J.-R.; Jung, J.-O.; Plum, P.; Loeser, H.; Lohneis, P.; Meuschke, M.; Pinto dos Santos, D.; Gebauer, F.; Quaas, A.; Bruns, C. J.; Walch, A.; Lawonn, K.; Popp, F. C.; Bozek, K.

2022-05-16 bioinformatics
10.1101/2022.05.13.491769 bioRxiv
Show abstract

BackgroundFast and accurate diagnostics are key for personalized medicine. Particularly in cancer, precise diagnosis is a prerequisite for targeted therapies which can prolong lives. In this work we focus on the automatic identification of gastroesophageal adenocarcinoma (GEA) patients that qualify for a personalized therapy targeting epidermal growth factor receptor 2 (HER2). We present a deep learning method for scoring microscopy images of GEA for the presence of HER2 overexpression. MethodsOur method is based on convolutional neural networks (CNNs) trained on a rich dataset of 1,602 patient samples and tested on an independent set of 307 patient samples. We incorporated an attention mechanism in the CNN architecture to identify the tissue regions in these patient cases which the network has detected as important for the prediction outcome. Our solution allows for direct automated detection of HER2 in immunohistochemistry-stained tissue slides without the need for manual assessment and additional costly in situ hybridization (ISH) tests. ResultsWe show accuracy of 0.94, precision of 0.97, and recall of 0.95. Importantly, our approach offers accurate predictions in cases that pathologists cannot resolve, requiring additional ISH testing. We confirmed our findings in an independent dataset collected in a different clinical center. ConclusionsWe demonstrate that our approach not only automates an important diagnostic process for GEA patients but also paves the way for the discovery of new morphological features that were previously unknown for GEA pathology.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.