Benchmarking artificial intelligence methods for end-to-end computational pathology
Laleh, N. G.; Muti, H. S.; Loeffler, C. M. L.; Echle, A.; Saldanha, O. L.; Mahmood, F.; Lu, M. Y.; Trautwein, C.; Langer, R.; Dislich, B.; Buelow, R. D.; Grabsch, H. I.; Brenner, H.; Chang-Claude, J.; Alwers, E.; Brinker, T. J.; Khader, F.; Truhn, D.; Gaisa, N. T.; Boor, P.; Hoffmeister, M.; Schulz, V.; Kather, J. N.
Show abstract
Artificial intelligence (AI) can extract subtle visual information from digitized histopathology slides and yield scientific insight on genotype-phenotype interactions as well as clinically actionable recommendations. Classical weakly supervised pipelines use an end-to-end approach with residual neural networks (ResNets), modern convolutional neural networks such as EfficientNet, or non-convolutional architectures such as vision transformers (ViT). In addition, multiple-instance learning (MIL) and clustering-constrained attention MIL (CLAM) are being used for pathology image analysis. However, it is unclear how these different approaches perform relative to each other. Here, we implement and systematically compare all five methods in six clinically relevant end-to-end prediction tasks using data from N=4848 patients with rigorous external validation. We show that histological tumor subtyping of renal cell carcinoma is an easy task which approaches successfully solved with an area under the receiver operating curve (AUROC) of above 0.9 without any significant differences between approaches. In contrast, we report significant performance differences for mutation prediction in colorectal, gastric and bladder cancer. Weakly supervised ResNet-and ViT-based workflows significantly outperformed other methods, in particular MIL and CLAM for mutation prediction. As a reason for this higher performance we identify the ability of ResNet and ViT to assign high prediction scores to highly informative image regions with plausible histopathological image features. We make all source codes publicly available at https://github.com/KatherLab/HIA, allowing easy application of all methods on any end-to-end problem in computational pathology.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 96%
- Multiple instance learning with pathology foundation models effectively predicts kidney disease diagnosis and clinical classification 94%
- PathProfiler: Automated Quality Assessment of Retrospective Histopathology Whole-Slide Image Cohorts by Artificial Intelligence, A Case Study for Prostate Cancer Research 94%
Similar papers in this journal
Similar papers in this journal
- A Spatial Attention Guided Deep Learning System for Prediction of Pathological Complete Response Using Breast Cancer Histopathology Images 94%
- Multi-Omic Graph Diagnosis (MOGDx) : A data integration tool to perform classification tasks for heterogeneous diseases 94%
- ELISL: Early-Late Integrated Synthetic Lethality Prediction in Cancer 93%
Similar papers in this journal
- Comparative Analysis of Pathology Foundation Models for Automated Detection of Tertiary Lymphoid Structures in H&E-Stained Digital Pathology Images 95%
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 94%
- Explainable Machine Learning for Preoperative Relapse Prediction in Molecularly Stratified Endometrial Cancer: A Single-Center Finnish Cohort Study 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.