Back

Self-Supervised Learning Can Distinguish Myelodysplastic Neoplasms from Clinical Mimics Using Bone Marrow Biopsies

Mehrtash, V.; Le, H.; Jafarzadeh, B.; Loghavi, S.; Garcia-Manero, G.; Tsirigos, A.; Park, C. Y.

2025-02-21 pathology
10.1101/2025.02.17.25322075 medRxiv
Show abstract

The diagnosis of myelodysplastic neoplasms (MDS) requires examination of the bone marrow for morphologic evidence of dysplasia. We sought to determine if a self-supervised learning (SSL) AI image analysis approach may be utilized to reliably distinguish MDS from its clinically relevant mimics using bone marrow biopsies (BMBx). Whole slide images (WSIs) of H&E- and reticulin-stained BMBx sections from 243 unique patients (89 MDS, 55 non-MDS cytopenic controls [NMCC], and 99 negative control [NC] cases) were segmented into tiles and analyzed. These tiles were then processed using the Barlow Twins SSL model to generate histomorphologic phenotype clusters (HPCs). Review of the HPCs revealed the clusters enriched in MDS captured known histopathologic features of MDS including hypercellularity, dysplastic and clustered megakaryocytes, increased immature hematopoietic cells, increased vascularity, fibrosis, and cell streaming patterns. Assessment of 95 MDS BMBx images from a second institution showed consistent HPC enrichment patterns, validating the robustness of the model. The trained ensemble model using H&E- and reticulin-stained slides distinguished MDS from NCs with an AUC of 0.89, and from age-matched, NMCCs with an AUC of 0.84. These findings demonstrate the potential of SSL approaches to capture diagnostically relevant morphologic patterns and to improve the reproducibility of MDS diagnosis.

Published in Blood Neoplasia · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.