Human Intuition vs. Computational Precision: Neurologists, Feature-based Models, and Deep Learning for Stroke Prognosis
Herzog, L.; Blindenbacher, N.; Globas, C.; Haeberlin, M. I.; Baumgartner, P.; Capecchi, F.; Inauen, C.; Sick, B.; Majoie, C. B.; van Zwam, W. H.; Wegener, S.
Show abstract
Background: Prognostication in large vessel occlusion (LVO) stroke remains challenging. Although several prognostic models exist, their comparison to clinician performance, human-model interaction, and specific sources of human bias remain poorly understood. Methods: Using pre-treatment clinical and CT data from the MR CLEAN trial (n=500), six neurologists predicted three-month modified Rankin Scale (mRS) scores for 40 patients, both unaided and assisted by a validated feature-based model (MR PREDICTS). Human performance was benchmarked against MR PREDICTS and a multimodal, interpretable deep learning (DL) approach using raw imaging data. We explicitly assessed neurologists? ability to estimate model-required imaging features and identified systematic human biases. Models were additionally validated in a larger MR CLEAN trial cohort (n=404). Results: For predicting the full mRS distribution, standalone models achieved good ordinal agreement (MR PREDICTS quadratic weighted kappa (QWK) 0.51 [0.24 to 0.70]; DL model 0.49 [0.25 to 0.67]), significantly outperforming unaided neurologists (QWK 0.27 [0.10, 0.42]). Neurologists showed systematic overoptimism, predicting lower mRS scores than observed. Furthermore, there was poor accuracy in extracting imaging features. Raters? ASPECTS predictions deviated by 3.4 points from the confirmed scores, and collateral score accuracy was 44.6%. However, for predicting binary mRS (0-2 vs. 3-6), accuracy was comparable between unaided neurologists (64.17% [55.42% to 72.92%]) and models (MR PREDICTS 67.50% [52.50% to 82.50%]; DL model 63.16% [47.37% to 78.95%]). Model-assistance modestly improved and harmonized neurologists? predictions (QWK 0.41 [0.22 to 0.55]; binary accuracy 68.75% [58.33% to 78.34%]. Model performance remained robust in the larger cohort. Conclusions: Multimodal prognostic models outperform clinicians in predicting the full range of mRS outcomes, while human error in imaging assessment and systematic optimism bias are primary drivers of prognostic inaccuracy. End-to-end DL models eliminate human-input variability and hold strong potential as an automated second opinion to support prognostication and decision-making in acute LVO stroke.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- End to end stroke triage using cerebrovascular morphology and machine learning 94%
- Automated Identification of Thrombectomy Amenable Vessel Occlusion on Computed Tomography Angiography using Deep Learning 93%
- Rich-Club organization: an important determinant of functional outcome after acute ischemic stroke 93%
Similar papers in this journal
- Blood-brain barrier leakage in the penumbra is associated with infarction on follow-up imaging in acute ischemic stroke 94%
- Effect of Time to Thrombolysis on Clinical Outcomes in Patients with Acute Ischemic Stroke Treated with Tenecteplase Compared to Alteplase: Analysis from the AcT Randomized Controlled Trial 93%
- Early recanalization among patients undergoing bridging therapy with tenecteplase or alteplase 93%
Similar papers in this journal
- A Clinical Neuroimaging Platform for Rapid, Automated Lesion Detection and Personalized Post-Stroke Outcome Prediction 93%
- Machine learning-based forecasting of daily acute ischemic stroke admissions using weather data 91%
- Multicenter Evaluation of Interpretable AI for Coronary Artery Disease Diagnosis from PET Biomarkers 90%
Similar papers in this journal
- Scaling behaviors of deep learning and linear algorithms for the prediction of stroke severity 95%
- Ground-truth validation of uni- and multivariate lesion inference approaches 90%
- Machine learning-based prediction of motor status in glioma patients using diffusion MRI metrics along the corticospinal tract 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.