Back

Explainable machine-learning model to classify culprit calcified carotid plaque in embolic stroke of undetermined source

Sakai, Y.; Kim, J.; Phi, H. Q.; Hu, A. C.; Balali, P.; Guggenberger, K. V.; Woo, J. H.; Bos, D.; Kasner, S. E.; Cucchiara, B.; Saba, L.; Huang, Z.; Haehn, D.; Song, J. W.

2024-10-30 radiology and imaging
10.1101/2024.10.25.24316081 medRxiv
Show abstract

BackgroundEmbolic stroke of undetermined source (ESUS) may be associated with carotid artery plaques with <50% stenosis. Plaque vulnerability is multifactorial, possibly related to intraplaque hemorrhage (IPH), lipid-rich-necrotic-core (LRNC), perivascular adipose tissue (PVAT), and calcification morphology. Machine-learning (ML) approaches in plaque classification are increasingly popular but often limited in clinical interpretability by black-box nature. We apply an explainable ML approach, using noncalcified plaque components and calcification features with SHapley Additive exPlanations (SHAP) framework to classify calcified carotid plaques as culprit/non-culprit. MethodsIn this retrospective cross-sectional study, patients with unilateral anterior circulation ESUS who underwent neck CT angiography and had calcific carotid plaque were analyzed. Calcification-level features were derived from manual segmentations. Plaque-level features were assessed by a neuroradiologist blinded to stroke-side and by semi-automated software. Calcifications/plaques were classified as culprit if ipsilateral to stroke-side. Eight baseline ML models were compared. Three CatBoost models were trained: Plaque-level, Calcification-level, and Combined. SHAP was incorporated to explain model decisions. Results70 patients yielded 116 calcific carotid plaques (60 ipsilateral to stroke; 270 calcifications (146 ipsilateral)). 17 plaque-level and 15 calcification-level features were extracted. Baseline CatBoost model outperformed other models. Combined model achieved test AUC 0.77 (95% CI: 0.59-0.92), accuracy 0.82 (95% CI: 0.71 - 0.91), mean cross-validation AUC 0.78. Plaque-level and calcification-level models performed lower (AUC 0.41 95% CI: 0.15-0.68, 0.60 95% CI 0.44-0.76). Combined model utilized five features: plaque thickness, IPH/LRNC volume ratio, PVAT volume, calcification minimum density, and total calcification volume over mean density ratio. Plaque thickness was most important feature based on SHAP values, with potential threshold at >2.6 mm. ConclusionsML model trained with noncalcified plaque and calcification features can classify culprit calcific carotid plaque with greater accuracy than models trained using only plaque-level or calcification-level features. Model using clinically interpretable features with SHAP framework provides explanations for its decisions and allows identification of potential thresholds for high-risk features. O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=110 SRC="FIGDIR/small/24316081v1_ufig1.gif" ALT="Figure 1"> View larger version (31K): org.highwire.dtl.DTLVardef@d7f3c7org.highwire.dtl.DTLVardef@1c5925aorg.highwire.dtl.DTLVardef@b5fedorg.highwire.dtl.DTLVardef@c6ff34_HPS_FORMAT_FIGEXP M_FIG O_FLOATNOGraphic AbstractC_FLOATNO Overall design of our study. C_FIG

Published in Journal of Neuroimaging · not in our set (fewer than 10 published preprints to learn from) · training set

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.