Domain specific models outperform large vision language models on cytomorphology tasks
Kukuljan, I.; Dasdelen, M. F.; Schaefer, J.; Buck, M.; Goetze, K.; Marr, C.
Show abstract
Large vision-language models (LVLMs) show impressive capabilities in image understanding across domains. However, their suitability for high-risk medical diagnostics remains unclear. We systematically evaluate four state-of-the-art LVLMs and three domain-specific models on key cytomorphological benchmarks: peripheral blood cell classification, morphology assessment, bone marrow cell classification, and cervical smear malignancy detection. Performance is assessed under zero-shot, few-shot, and fine-tuned conditions. LVLMs underperform significantly: the best LVLM achieves a zero-shot F1 score of 0.057 {+/-} 0.008 for malignancy detection--near random (0.039)--and only 0.15 {+/-} 0.01 in few-shot. In contrast, domain-specific models reach up to 0.83 in accuracy. Even after fine-tuning, a dedicated hematology model outperforms GPT-4o. While LVLMs offer explainability via text, we find the visual-language grounding unreliable, and the morphological features mention by the model often do not match the single cell properties. Our findings suggest that LVLMs require substantial improvements before use in high-stakes diagnostic settings. Key findingsO_LILVLMs perform poorly on cytomorphology tasks, often near chance level and far below domain-specific models. C_LIO_LIEven after fine-tuning, LVLMs lag behind domain-specific models. C_LIO_LIWhile LVLMs provide textual justifications, these often reflect generic descriptions rather than image-specific morphological features. C_LI
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Knowledge transfer to enhance the performance of deep learning models for automated classification of B-cell neoplasms 94%
- Federated Learning for multi-omics: a performance evaluation in Parkinson's disease 93%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Annotation-Free Deep Learning for Predicting Gene Mutations from Whole Slide Images of Acute Myeloid Leukemia 95%
- A Deep Learning Model for Molecular Label Transfer that Enables Cancer Cell Identification from Histopathology Images 94%
- Generalizing AI-driven Assessment of Immunohistochemistry across Immunostains and Cancer Types: A Universal Immunohistochemistry Analyzer 93%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.