The Expertise Paradox: Who Benefits from LLM-Assisted Brain MRI Differential Diagnosis?
Schramm, S.; Le Guellec, B.; Topka, M.; Svec, M.; Backhaus, P.; Eisenkolb, V. M.; Riedel, E. O.; Beyrle, M.; Platzek, P.-S.; Ramschütz, C.; Paprottka, K. J.; Renz, M.; Bodden, J.; Kirschke, J. S.; Ziegelmeyer, S.; Busch, F.; Makowski, M. R.; Adams, L. C.; Bressem, K. K.; Hedderich, D. M.; Wiestler, B.; Kim, S. H.
Show abstract
PurposeTo evaluate how reader experience influences the diagnostic benefit from LLM assistance in brain MRI differential diagnosis. Materials and MethodsNeuroradiologists (n = 4), radiology residents (n = 4), and neurology/neurosurgery residents (n = 4) were recruited. A dataset of complex brain MRI cases was curated from the local imaging database (n = 40). For each case, readers provided a textual description of the main imaging finding and their top three differential diagnoses ("Unassisted"). Three state-of-the-art large language models (GPT-4.1, Gemini 2.5 Pro, DeepSeek-R1) were prompted to generate top-three differentials based on the clinical case description and reader-specific findings. Readers then revised their differential diagnoses after reviewing GPT-4.1 suggestions ("Assisted"). To evaluate the association between reader experience and diagnostic benefit, a cumulative link mixed model (CLMM) was fitted, with change in diagnostic result as ordinal outcome, reader experience as predictor, and random intercepts for rater and case. ResultsLLM-generated differential diagnoses achieved the highest top-3 accuracy when provided with image descriptions from neuroradiologists (top-3: 78.8-83.8%), followed by radiology residents (top-3: 71.8-77.6%), and neurology/neurosurgery residents (top-3: 62.6-64.5%). In contrast, mean relative gains in top-3 accuracy through LLM assistance diminished with increasing experience, with +19.2% for neurology/neurosurgery residents (from 43.2% to 62.6%), +14.7% for radiology residents (from 59.6% to 74.4%), and +4.4% for neuroradiologists (from 83.1% to 87.5%). The CLMM demonstrated a significant negative association between reader experience and diagnostic benefit from LLM assistance ({beta} = -0.10, p = 0.005). ConclusionWith increasing reader experience, absolute diagnostic LLM performance with reader-generated input improved, while relative diagnostic gains through LLM assistance paradoxically diminished. Our findings call attention to the divergence between standalone LLM performance and clinically relevant reader benefit, and emphasize the need to account for human-AI interaction in this context.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Evaluating Large Language Model-Generated Brain MRI Protocols: Performance of GPT4o, o3-mini, DeepSeek-R1 and Qwen2.5-72B 98%
- Assessing GPT-4 Multimodal Performance in Radiological Image Analysis 95%
- Impact of Non-Contrast Enhanced Imaging Input Sequences on the Generation of Virtual Contrast-Enhanced Breast MRI Scans using Neural Networks 91%
Similar papers in this journal
- tbiExtractor: A framework for Extracting Traumatic Brain Injury Common Data Elements from Radiology Reports 92%
- Image-localized Biopsy Mapping of Brain Tumor Heterogeneity: A Single-Center Study Protocol 92%
- Classification performance bias between training and test sets in a limited mammography dataset 91%
Similar papers in this journal
- Content-based image retrieval assists radiologists in diagnosing eye and orbital mass lesions in MRI 95%
- Inconsistency of AI in Intracranial Aneurysm Detection with Varying Dose and Image Reconstruction 93%
- Deep neural networks allow expert-level brain meningioma detection, segmentation and improvement of current clinical practice 93%
Similar papers in this journal
- A Clinical Neuroimaging Platform for Rapid, Automated Lesion Detection and Personalized Post-Stroke Outcome Prediction 91%
- EchoGraph System for Automated Quality Assessment of Echocardiography Reports 90%
- From Tool to Teammate: A Randomized Controlled Trial of Clinician-AI Collaborative Workflows for Diagnosis 90%
Similar papers in this journal
- Automated Tumor Segmentation and Brain Tissue Extraction from Multiparametric MRI of Pediatric Brain Tumors: A Multi-Institutional Study 93%
- Pediatric brain tumor classification using deep learning on MR-images with age fusion 91%
- Early prognostication of overall survival for pediatric diffuse midline gliomas using MRI radiomics and machine learning 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.