Back

Biomarker Fidelity Score - A Quantitative Framework for Individual-Level Validation of Explainability Methods in 3D Alzheimer's Disease MRI Classification

Lepcha, D. C.; Ali, A.; Martin, S. A.; Syed-Abdul, S.

2026-08-20 neuroscience
10.64898/2026.08.15.744687 bioRxiv
Show abstract

Explainability methods applied to deep learning models for Alzheimer's disease neuroimaging produce attribution maps that vary substantially across methods and architectures, yet no validated quantitative framework exists for determining which method most faithfully localises attribution signal within established AD biomarker anatomy at the individual subject level. Existing validation approaches rely on group-level comparisons or qualitative visual inspection, leaving individual-level biomarker alignment uncharacterised. We introduce the Biomarker Fidelity Score (BFS), a quantitative tool measuring spatial overlap between individual-level 3D explainability attention maps and atlas-registered AD-relevant neuroimaging ROIs across thirteen anatomically defined structures including hippocampus, entorhinal cortex, amygdala, and parahippocampal gyrus. Five explainability methods (GradCAM++, Integrated Gradients, DeepSHAP, LRP, ScoreCAM) were benchmarked across three volumetric architectures (3D ResNet-18, DenseNet-121, Swin-UNETR) on 327 balanced ADNI-3 subjects. Integrated Gradients achieved the highest BFS across all architectures while GradCAM++ consistently showed the lowest biomarker alignment (all p<0.001, Friedman test). The complete BFS pipeline replicated these rankings without retraining on 207 independent OASIS-3 subjects, with maximum absolute difference of 0.0005 across all fifteen method-architecture combinations and Spearman rank correlation of 0.964 between cohort rankings. By offering an externally validated, individual-level, biomarker-grounded quantitative standard, BFS equips clinicians and AI developers with practical guidance for selecting trustworthy explainability methods in AD neuroimaging.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.