Benchmarking ten frontier large language models on 1,477 board style multiple choice questions in hematology
Radoynova, M.; Benouis, M.; schulze, f.; Winter, S.; Bornhauser, M.; Middeke, J. M.; Eckardt, J.-N.
Show abstract
Large Language Models (LLMs) are increasingly used by clinicians and patients for medical queries, yet their accuracy and safety at the specialist level in hematology remain insufficiently characterised. We benchmarked ten frontier proprietary and open-weight LLMs across two generations on 1,477 board-style hematology multiple-choice questions (MCQs) derived from five educational datasets spanning nine disease areas and six clinical skill domains, including text-only and multimodal case vignettes. Claude Opus 5 had the highest mean accuracy (92.7% text, 76.9% multimodal), followed closely by Gemini-3.1 Pro (91.4% and 78.7%), Gemini-3.6 Flash (91.0% and 74.8%) and GPT-5.6 Sol (89.9% and 76.7%). Accuracy significantly correlated with model size both for text-only and multimodal MCQs. Between model generations, the largest improvements in accuracy were seen for open-weight models whereas proprietary models showed only marginal gains. In error analysis, top-performing models exhibited highly concordant failure patterns, suggesting shared limitations on challenging cases. Frontier LLMs exhibit substantial specialist hematology knowledge across diverse subspecialist domains and clinical skill sets. Yet, despite high accuracy on board-style questions in hematology, continuous expert-on-the-loop output monitoring is paramount.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 93%
- Interpretable Fine-tuned Large Language Models Facilitate Making Genetic Test Decisions for Rare Diseases 92%
- Evaluating large language model workflows in clinical decision support: referral, triage, and diagnosis 91%
Similar papers in this journal
- Automated abstraction of clinical parameters of multiple myeloma from real-world clinical notes using large language models 91%
- Therapeutic Monoclonal Antibodies Repurposing in Oncology via IMGT/mAb-KG Embeddings 90%
- ARDSFlag: An NLP/Machine Learning Algorithm to Visualize and Detect High-Probability ARDS Admissions Independent of Provider Recognition and Billing Codes 90%
Similar papers in this journal
- Measuring COVID-19 and Influenza in the Real World via Person-Generated Health Data 89%
- Knowledge transfer to enhance the performance of deep learning models for automated classification of B-cell neoplasms 89%
- Structuring clinical text with AI: old vs. new natural language processing techniques evaluated on eight common cardiovascular diseases 89%
Similar papers in this journal
- Exploring the role of Large Language Models (LLMs) in hematology: a systematic review of applications, benefits, and limitations 91%
- COVID symptoms, testing, shielding impact on patient reported outcomes and early vaccine responses in individuals with multiple myeloma 85%
- A validation study of the identification of haemophagocytic lymphohistiocytosis in England using population-based health data 84%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.