Back

On the assessment of deep-learning based super-resolution in small datasets of human brain MRI scans

Loeffen, D. W. M.; Rijpma, A.; Bartels, R. H. M. A.; Vinke, R. S.

2026-02-17 radiology and imaging
10.64898/2026.02.16.26346392 medRxiv
Show abstract

Deep-learning based super-resolution has shown promise for enhancing the spatial resolution of brain magnetic resonance images, which may help visualize small anatomical structures more clearly. However, when only limited training data are available, it remains uncertain which model assessment method provides the most reliable estimate of out-of-sample performance. In this study, three widely used assessment strategies (three-way holdout, k-fold cross-validation, and nested cross-validation) were compared for evaluating the performance of such models in small datasets. Across 30 iterations, we randomly selected subsets of 20 T2-weighted images from the 1,113 scans of the Human Connectome Project. Each subset was used to train a model and estimate performance using the three methods. The ground truth error was computed from the remaining images. The assessment error is the difference between the estimated error and the ground truth error. The median assessment errors were 0.11,- 0.13, - 0.32 for three-way holdout, k-fold cross-validation, and nested cross-validation, respectively, with the cross-validation methods showing considerably smaller dispersions. Nested cross-validation selected fewer epochs, indicating more conservative model selection, but required substantially greater computational time, over three times longer than three-way holdout and more than twenty times longer than k-fold cross-validation. Our findings suggest that k-fold cross-validation offers the most favourable balance between accuracy, stability, and computational feasibility in small datasets. Further research is needed to determine how model complexity, dataset size, and the number of cross-validation folds influence assessment accuracy.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.