Back

Modern Convolutional Design Improves Uterine MRI Segmentation, while nnU-Net Remains Most Robust Across Heterogeneous Datasets

Di Giovanni, D. A.; Takada, A.; McNabb, E.; Dana, J.; Yokota, H.; Tsuboyama, T.; Zakarian, R.; Vallieres, M.; Tsui, J. M. G.; Reinhold, C.

2026-07-24 radiology and imaging
10.64898/2026.07.22.26358583 medRxiv
Show abstract

Purpose: To evaluate how segmentation architecture and dataset-adaptive configuration influence uterine MRI segmentation across heterogeneous benign and malignant tasks. Methods: U-Net, Swin-UNETR, and MedNeXt were compared with nnU-Net as a self-configuring reference across T2-weighted MRI datasets: public multiclass UMD anatomy/fibroid segmentation (n=300), institutional endometrial cancer tumor segmentation (n=206), and institutional uterine mass lesion segmentation (n=234). A relabeled external UMD-style cohort (n=12) assessed domain shift. Models used fixed partitions, fold ensembling, Dice, HD95, ASSD, volume error, and paired bootstrap comparisons with Holm correction. Results: MedNeXt was the strongest manually controlled architecture. nnU-Net achieved the highest performance on all internal datasets and external testing. Macro-Dice reached 0.761, 0.746, and 0.814 for nnU-Net on UMD, endometrial cancer, and uterine mass datasets, respectively, versus 0.722, 0.726, and 0.789 for MedNeXt. The nnU-Net-MedNeXt gap was largest for multiclass UMD segmentation and smaller in binary tasks. External testing degraded all models; nnU-Net remained highest (0.542), followed by MedNeXt (0.490), U-Net (0.396), and Swin-UNETR (0.287). Conclusions: Uterine MRI segmentation performance depended on task, architecture, and evaluation domain. MedNeXt supported modern convolutional design as a strong manual baseline, but nnU-Net remained the most robust overall, emphasizing the importance of dataset-adaptive configuration and external validation.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.