Back

Unsupervised Tissue Concepts for Explainable Sarcoma Subtype Prediction from H&E

Bisson, T.; Ingram, D.; Singh, S.; Li, A.; Flynn, S.; Wang, W.-L.; Kim, A. E.; Bridge, C. P.; Demicco, E. G.; Sorrentino, A.; Jiang, S.; Hung, Y. P.; Lazar, A. J.; Iafrate, A. J.

2026-05-20 pathology
10.64898/2026.05.15.26353333 medRxiv
Show abstract

Soft tissue sarcomas are a rare, heterogeneous group of tumors whose diagnosis remains challenging because of overlapping morphology and limited access to sarcoma-specialized pathologists. Although pathology foundation models have shown promise in computational pathology, their clinical translation remains limited by insufficient interpretability, particularly in diagnostically complex settings such as sarcoma diagnosis. Here, we developed and evaluated an H&E-based AI framework for sarcoma subtype classification that focused on explanability. Using the CONCH v1.5 foundation model, we computed embeddings from a tissue microarray cohort of 2,545 cases spanning 19 sarcoma subtypes and trained an attention-based multiple-instance learning model that achieved a balanced accuracy of 77.38% (SD 1.88). To move explainability beyond attention-based localization, we trained a sparse autoencoder on patch-level embeddings to learn 768 recurring visual concepts. 90 high-activation concepts were reviewed by three senior pathologists and curated into morphologically meaningful and non-meaningful categories, yielding a semantic dictionary of 41 diagnostically relevant tissue concepts. We then trained a linear attention-based model on the 768-concept vectors, which retained much of the performance of the raw embedding-based ABMIL model, achieving a balanced accuracy of 73.74% (SD 1.30). When restricting the linear model to pathologist-curated morphologic concepts only, balanced accuracy further decreased to 67.04% (SD 1.27), suggesting that the residual performance gain in the full concept model was driven by inconsistent, technical, or diagnostically irrelevant concepts. Concept-level explanations of the curated linear attention-based model aligned with known sarcoma morphology, including lipogenic, myxoid, spindle-cell, pleomorphic, vascular, small round blue cell, and matrix-forming patterns, and reproduced patterns of diagnostic overlap observed in human sarcoma pathology. Together, these results show that H&E-based foundation-model representations capture meaningful diagnostic structure within the known limitations of H&E in sarcoma diagnostics, but that their clinical value depends on whether this structure can be made interpretable to pathologists. Sparse autoencoder-derived concepts can address this critical gap by converting embedding-level signal into recurring morphologic patterns that pathologists can review and name, providing the foundation to link these patterns to subtype predictions. In doing so, this approach turns concept discovery into a practical form of diagnostic explanation, while also revealing where model performance is supported by recognizable histopathology and where it relies on diagnostically irrelevant or inconsistent visual patterns.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

1
Modern Pathology
22 papers in training set
Top 0.1%
27.0%
2
npj Digital Medicine
118 papers in training set
Top 0.5%
12.7%
3
npj Precision Oncology
53 papers in training set
Top 0.1%
8.0%
4
Communications Medicine
113 papers in training set
Top 0.3%
6.4%
50% of probability mass above
5
Nature Communications
5641 papers in training set
Top 32%
4.1%
6
Nature Machine Intelligence
70 papers in training set
Top 0.8%
3.3%
7
Journal of Pathology Informatics
15 papers in training set
Top 0.1%
3.3%
8
Nature Medicine
125 papers in training set
Top 0.8%
2.8%
9
PLOS Computational Biology
1863 papers in training set
Top 11%
2.5%
10
Scientific Reports
3612 papers in training set
Top 46%
2.2%
11
Medical Image Analysis
35 papers in training set
Top 0.4%
1.9%
12
The American Journal of Pathology
32 papers in training set
Top 0.3%
1.7%
13
Cancers
213 papers in training set
Top 3%
1.5%
14
eBioMedicine
183 papers in training set
Top 4%
1.1%
15
Science Translational Medicine
127 papers in training set
Top 2%
1.1%
16
Science Advances
1243 papers in training set
Top 24%
1.1%
17
The Journal of Pathology
26 papers in training set
Top 0.5%
1.1%
18
Cell Reports Medicine
153 papers in training set
Top 4%
1.1%
19
Clinical Cancer Research
64 papers in training set
Top 2%
1.0%
20
Advanced Science
286 papers in training set
Top 9%
0.9%
21
PLOS ONE
5266 papers in training set
Top 60%
0.9%
22
Journal of Medical Imaging
11 papers in training set
Top 0.4%
0.9%
23
Breast Cancer Research
36 papers in training set
Top 0.5%
0.9%
24
BMC Cancer
67 papers in training set
Top 2%
0.9%
25
JAMIA Open
42 papers in training set
Top 2%
0.6%
26
Cancer Discovery
66 papers in training set
Top 2%
0.6%
27
Biology Methods and Protocols
61 papers in training set
Top 3%
0.6%
28
Cancer Research Communications
51 papers in training set
Top 2%
0.6%
29
eLife
5828 papers in training set
Top 68%
0.6%