Back

Counterfactual Diffusion Models for Mechanistic Explainability of Artificial Intelligence Models in Pathology

Zigutyte, L.; Lenz, T.; Han, T.; Hewitt, K. J.; Reitsam, N. G.; Foersch, S.; Carrero, Z.; Unger, M.; Pearson, A. T.; Truhn, D.; Kather, J. N.

2025-01-08 bioinformatics
10.1101/2024.10.29.620913 bioRxiv
Show abstract

Deep learning can extract predictive and prognostic biomarkers from histopathology whole slide images. However, explainable artificial intelligence approaches widely used in digital pathology, such as attention heatmaps and class activation mapping, offer only limited interpretability regarding the features captured by classifiers. Here, we present MoPaDi (Morphing histoPathology Diffusion), a framework for generating counterfactual explanations for histopathology images that reveal which morphological or style features drive classifier predictions. MoPaDi combines diffusion autoencoders with task-specific multiple instance learning classifiers to manipulate images and flip predictions by modifying relevant features. We evaluated the framework on multiple datasets spanning colorectal, breast, liver, and lung cancers, including tissue type, cancer subtype, and biomarker (microsatellite instability) classification tasks. We assessed counterfactual explanations through quantitative analyses, pathologists evaluations, and independent foundation model-based classifiers. We found that MoPaDi was able to generate realistic counterfactual histopathology images, enabling pathologists to identify morphological features associated with the change in model predictions. Unlike conventional reviews of highly attended regions typical in digital pathology, MoPaDi explanations enabled pathologists to directly identify morphological features driving the classifiers predictions from a limited number of top-contributing tiles. Consistent with the literature, our biomarker classifier associated high microsatellite instability with mucinous differentiation, glandular patterns, and lymphocytic infiltration. Furthermore, MoPaDi revealed that changes in classifier predictions were mainly driven by morphological alterations rather than staining differences. Overall, MoPaDi is a practical framework for counterfactual explanations in computational pathology that reveals model-specific drivers of classification and increases trust in deep learning models.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.