Three multimodal large language models fail at clinically actionable breast pathology in three different directions
Kang, Y.-J.; Jun, S.-Y.; Kim, S.
Show abstract
Background. Breast cancer treatment depends on histopathological features, such as grade and receptor-defined subtype; however, specialist pathologist access is constrained when the workforce is limited. Commercial multimodal large language models (MLLMs) accept hematoxylin and eosin (H&E) image tiles through paid interfaces without local hardware or fine-tuning. However, prior pathology evaluations addressed only coarse tasks. Whether they reach treatment-determining accuracy and whether vendors agree remain unclear. Methods. We aimed to evaluate three vendor-designated flagship MLLMs (Claude Sonnet 4.6, Gemini 2.5 Pro, GPT-5.5) in 427 invasive breast cancer cases. Each case went to all three with identical H&E tiles and prompts, and the subtype was inferred in the second call. The reference was an institutional sign-out report of an immunohistochemistry-derived subtype. We calculated the concordance, sensitivity, specificity, Cohen's kappa, and pairwise McNemar and Bowker tests. Findings. Claude ranked highest by raw histologic-type concordance but lowest by kappa, classifying all 23 lobular and seven micropapillary carcinomas as invasive breast carcinoma of no special type. The models anchored the Nottingham grade to three modal grades. None of the models reliably identified human epidermal growth factor receptor 2-positive disease. The failure direction was vendor-specific: Claude and GPT-5.5 were under-detected, whereas Gemini was over-called. Twelve prompt variants (4,056 calls) did not recover sensitivity. Interpretation. No current commercial MLLM reaches deployment-ready accuracy for any treatment-determining feature of breast pathology. As each vendor fails in its own fixed direction, changing vendors alters the type of error rather than removing it; therefore, the value of these models is assistive rather than autonomous. At USD 0.20-0.50 per case, they may serve as supervised draft generators that leave the diagnosis with the pathologist.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Genomic Characterization of Lung Cancer in Never-Smokers Using Deep Learning 94%
- Attention-based whole-slide image compression achieves pathologist-level pre-screening of multi-organ routine histopathology biopsies 93%
- Tissue contamination challenges the credibility of machine learning models in real world digital pathology 93%
Similar papers in this journal
- Relation of Quantitative Histologic and Radiologic Breast Tissue Composition Metrics with Invasive Breast Cancer Risk 92%
- Deep Learning Image Analysis of Benign Breast Disease to Identify Subsequent Risk of Breast Cancer 92%
- Tumor localization strategies of multi-cancer early detection tests: a quantitative assessment 92%
Similar papers in this journal
- Automated quantification of Ki-67 expression in breast cancer from H&E-stained slides using a transformer-based regression model 95%
- Development and prognostic validation of a three-level NHG-like deep learning-based model for histological grading of breast cancer 93%
- Effect of testosterone therapy on breast tissue composition and mammographic breast density in trans masculine individuals 92%
Similar papers in this journal
- RNA Sequencing-Based Single Sample Predictors of Molecular Subtype and Risk of Recurrence for Clinical Assessment of Early-Stage Breast Cancer 94%
- Unmasking the tissue microecology of ductal carcinoma in situ with deep learning 93%
- An updated PREDICT breast cancer prognostic model including the benefits and harms of radiotherapy 92%
Similar papers in this journal
- Development and validation of AI-based pre-screening of large bowel biopsies 93%
- Novel deep learning algorithm predicts the status of molecular pathways and key mutations in colorectal cancer from routine histology images 92%
- Multicenter Validation of a Machine Learning Algorithm for Diagnosing Pediatric Patients with Multisystem Inflammatory Syndrome and Kawasaki Disease 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.