Vision Language Models Fail to Reliably Detect Acute Myeloid Leukemia in Bone Marrow Smears
Schulze, F.; Loeffler, C.; Radoynova, M.; Winter, S.; Roellig, C.; Sockel, K.; Kroschinsky, F.; Bornhaeuser, M.; Middeke, J. M.; Kather, J. N.; Eckardt, J.-N.; Ghaffari Laleh, N.
Show abstract
Hematologic diagnostics and especially cytomorphologic assessment are time-intensive and require high levels of expertise. Vision Language Models (VLM) show promise in medical image analysis in radiology and histopathology, while an evaluation on detecting acute myeloid leukemia (AML) is lacking. Our goal was to evaluate three Vision Language Models regarding their diagnostic accuracy and safety in clinical decision support in detecting AML from digitized bone marrow smears (BMS). Whole slide images were obtained from bone marrow smears of 50 AML patients and 50 bone marrow donors. Ten representative fields of view per sample were extracted manually. Three VLMs were used, two of which are considered generalist models (Qwen3.5-397B-A17B-FP8, GLM-4.6V-FP8), while the other one is a medically adapted model (Medgemma-27b-it). All models performed zero-shot analysis using two prompting strategies: First, a context-rich prompt requesting reporting of WHO/FAB diagnostic criteria in a structured manner, and secondly a minimal prompt without specific hematologic context. Overall diagnostic accuracy was poor for all models as they exhibited the overwhelming tendency to classify most samples as leukemic: With context-rich prompts, GLM4.6 identified 90% of leukemic samples while also labeling 92% of bone marrow donors as AML. The medical specialist model MedGemma-27b showed similar failure, misclassifying 86% of healthy donors and correctly detecting AML in only 66% of cases. Qwen3.5 performed best under detailed prompting, achieving a specificity of 0.26 and accuracy of 0.51. Accuracy of all models improved with context-free prompts (accuracies range 0.47-0.79), yet they still lacked the ability to correctly distinguish between leukemia and healthy bone marrow. Qwen3.5 was the only model to maintain meaningful specificity (0.64) and correctly identified 94% of AML, yielding an overall accuracy of 0.79. Morphologic feature-level agreement with human expert reports was poor across all models, indicating poor recognition of cell-level morphologies. This failure is likely driven by the fact that pathology imaging archives are vastly scraped during model training while hematological samples are not as widely available and therefore, hematology is an out-of-bounds use-case for these models, rendering them currently unsuitable for clinical decision support in hematology.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Hematologist-level classification of mature B-cell neoplasm using deep learning on multiparameter flow cytometry data 92%
- Classification of human white blood cells using machine learning for stain-free imaging flow cytometry 92%
- DeepIFC: virtual fluorescent labeling of blood cells in imaging flow cytometry data with deep learning 90%
Similar papers in this journal
- Image-based Explainable Artificial Intelligence Accurately Identifies Myelodysplastic Neoplasms Beyond Conventional Signs of Dysplasia 95%
- Annotation-Free Deep Learning for Predicting Gene Mutations from Whole Slide Images of Acute Myeloid Leukemia 92%
- Multimodal Spatial Proteomic Profiling in Acute Myeloid Leukemia 87%
Similar papers in this journal
- Differential response to cytotoxic therapy explains treatment dynamics of AML patients: insights from a mathematical modelling approach 86%
- Mitigating Temozolomide Resistance in Glioblastoma via DNA Damage-Repair Inhibition 84%
- Mathematical deconvolution of CAR T-cell proliferation and exhaustion from real-time killing assay data 84%
Similar papers in this journal
- Enhancing Semantic Segmentation in Chest X-Ray Images through Image Preprocessing: ps-KDE for Pixel-wise Substitution by Kernel Density Estimation 90%
- ai-corona : Radiologist-Assistant Deep Learning Framework for COVID-19 Diagnosis in Chest CT Scans 89%
- Analytical validation and performance characteristics of a 48-gene next-generation sequencing panel for detecting potentially actionable genomic alterations in myeloid neoplasms 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.