Benchmarking Generalizability in Deep Learning-Based White Matter Tract Segmentation
Kwon, J.; Amorosino, G.; Pestilli, F.
Show abstract
Summary paragraphWhite matter tracts (WMTs) are the brains structural foundation for information transfer, underlying essential cognitive and behavioral functions. While diffusion MRI and tractography enable non-invasive mapping of these pathways, automated segmentation often lacks generalizability across diverse data sources. We conducted a systematic, cross-dataset evaluation of four state-of-the-art deep learning architectures, benchmarking their performance across independent datasets with varying acquisition protocols and populations. CNN-based models such as TractSeg achieved the highest within-domain accuracy, but performance dropped sharply under domain shift, most severely when we applied adult-trained models to pediatric data. To address this degradation, we introduce Ensemble White Matter Tract Segmentation (EWMTS), which combines complementary models to partially recover accuracy under domain shift, although performance still falls short of within-domain levels. By openly releasing this benchmark and a reproducible processing pipeline, we provide the neuroimaging community with a framework to develop and benchmark segmentation models across the heterogeneity of real-world neuroimaging data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Automated joint skull-stripping and segmentation with Multi-Task U-Net in large mouse brain MRI databases 95%
- Insights from the IronTract challenge: optimal methods for mapping brain pathways from multi-shell diffusion MRI 95%
- Tensor Image Registration Library: Automated Deformable Registration of Stand-Alone Histology Images to Whole-Brain Post-Mortem MRI Data 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.