Board-Level Performance of Leading Open-Weight Vision-Language Models on the Japanese Diagnostic Radiology Board Examination: Reasoning, Image-Input, and Language Effects
Sonoda, Y.; Yamagishi, Y.; Hirano, Y.; Miki, S.; Nakao, T.; Hanaoka, S.; Nomura, Y.; Hamada, A.; Kanemaru, N.; Miyo, R.; Takahashi, M. M.; Hosoi, R.; Yoshikawa, T.; Abe, O.
Show abstract
Purpose: To evaluate the latest open-weight vision-language models (VLMs) on the Japanese Diagnostic Radiology Board Examination (JDRBE), assessing overall accuracy and the effects of image input, reasoning, and language. Materials and Methods: In this retrospective study, 29 open-weight VLMs from 13 developers, released in or after January 2025, were evaluated on 327 image-bearing questions from four years of the JDRBE, a non-public benchmark with low risk of data leakage. Each question was answered by each model with and without the image(s), under three language conditions and with reasoning enabled and disabled. Accuracy was the primary outcome, and within-model differences were tested with paired bootstrap confidence intervals and sign-flip permutation tests with Benjamini-Hochberg correction. Results: In the Japanese condition with image input and reasoning, the leading models reached 73.7% (gemma-4-31B-it), 73.1% (Qwen3.5-397B-A17B), and 72.1% (Kimi-K2.6). On the 2025 subset, these three models (74.1%-75.5%) scored above the mean accuracy of five newly board-certified radiologists who passed the 2025 examination (72%; range, 65%-83%). Accuracy broadly scaled with model size, although compact gemma-4-31B-it matched larger models. Enabling reasoning improved accuracy in nearly all models and the contribution of image input was larger when reasoning was enabled, particularly in higher-performing models. English prompts generally outperformed Japanese prompts. Conclusion: Several open-weight VLMs, without medical adaptation, performed at or above the mean of newly board-certified radiologists on the JDRBE, with both model size and reasoning contributing. The highest Japanese-language accuracy came from a compact model suitable for parameter-efficient fine-tuning and serving on a single graphics processing unit.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Performance of Generative Pretrained Transformer on the National Medical Licensing Examination in Japan 93%
- Development and Validation of a Deep Learning Model for Detecting Signs of Tuberculosis on Chest Radiographs among US-bound Immigrants and Refugees 93%
- Collaborative intelligence in AI: Evaluating the performance of a council of AIs on the USMLE 92%
Similar papers in this journal
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 92%
- Machine Learning Generalizability Across Healthcare Settings: Insights from multi-site COVID-19 screening 91%
- The clinician-AI interface: intended use and explainability in FDA-cleared AI devices for medical image interpretation 90%
Similar papers in this journal
- Large Language Models Improve the Identification of Emergency Department Visits for Symptomatic Kidney Stones 93%
- Reproducible And Clinically Translatable Deep Neural Networks For Cervical Screening 92%
- Generating synthetic data in digital pathology through diffusion models: a multifaceted approach to evaluation 92%
Similar papers in this journal
- Leveraging synthetic data produced from museum specimens to train adaptable species classification models 91%
- The relationship between confidence and gaze-at-nothing oculomotor dynamics during decision-making 90%
- Classification performance bias between training and test sets in a limited mammography dataset 89%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.