Impact of simulated MRI artifacts on deep learning-based brain age prediction
Hendriks, J.; Jansen, M. G.; Joules, R.; Pena-Nogales, O.; Elsen, F.; Povolotskaya, A.; Dijsselhof, M. B. J.; Rodrigues, P. R.; Barkhof, F.; Schrantee, A.; Mutsaerts, H.
Show abstract
Brain age is a promising biomarker for detecting atypical and pathological brain aging, but its accuracy and reliability depend critically on MRI quality. The impact of common MR image degradations such as motion, ghosting, blurring, and noise on brain age predictions remains unclear. In this study, we systematically assessed the effects of four simulated MRI artifact types, across ten severity levels, on brain age prediction using three widely used deep learning-based algorithms (Pyment, MIDI, MCCQR), in high-quality T1-weighted images of healthy adults (age range 18-85, 54% female). Artifact severity levels (1-10) were generated using a power-function mapping of TorchIO simulation parameters calibrated to the full PondrAI QC visual rating scale (from perfect to severely degraded image quality). Linear mixed-effects models with predicted brain age as dependent variable revealed a significant interaction between algorithm, artifact type, and severity (p<0.001), indicating algorithm-specific sensitivity to artifacts. In artifact-free scans, mean absolute error (MAE) was 4.6 years for MCCQR, 7.1 years for Pyment, and 9.1 years for MIDI. At severity level 10, MAE increased with up to 110% for Pyment, 112% for MCCQR, and 16% for MIDI (motion); and with up to 75% for Pyment, 135% for MCCQR, and 34% for MIDI (ghosting). Blurring had minimal impact at low-moderate levels, but at maximum severity MAE increased by 26% for Pyment and 137% for MCCQR, while MIDI remained largely stable. Noise minimally affected Pyment and MCCQR (MAE increases [≤]20%), but led to larger declines for MIDI (MAE increase 35%). The vulnerability of different algorithms highlights that training data, preprocessing strategies and underlying architectures influence robustness, emphasizing that artifact sensitivity is a key consideration when interpreting brain-age as a biomarker. Our results emphasize the need for artifact-aware evaluation and mitigation strategies when algorithms such as brain age are used in clinical research.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Internally-consistent and fully-unbiased multimodal MRI brain template construction from UK Biobank: Oxford-MM 96%
- Normative trajectories of R1, R2* and magnetic susceptibility in basal ganglia on healthy ageing 96%
- Predicting brain age across the adult lifespan with spontaneous oscillations and functional coupling in resting brain networks captured with magnetoencephalography 95%
Similar papers in this journal
- High spatial overlap but diverging age-related trajectories of cortical MRI markers aiming to represent intracortical myelin and microstructure 97%
- Brain-Age Prediction: Systematic Evaluation of Site Effects, and Sample Age Range and Size 96%
- Brain Age Prediction: Deep Models Need a Hand to Generalize 96%
Similar papers in this journal
Similar papers in this journal
- ComBat Harmonization: Empirical Bayes versus Fully Bayes Approaches 95%
- Cortical thickness and grey-matter volume anomaly detection in individual MRI scans: Comparison of two methods 95%
- Reliable longitudinal brain age prediction in stroke patients: Associations with cognitive function and response to cognitive training 95%
Similar papers in this journal
- Age-related differences in fMRI subsequent memory effects are directly linked to local grey matter volume differences 94%
- GABA levels in ventral visual cortex decline with age and are associated with neural distinctiveness 94%
- It is the Locus Coeruleus! Or... is it? : A proposition for analyses and reporting standards for structural and functional magnetic resonance imaging of the noradrenergic Locus Coeruleus 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.