Back

Benchmarking the robustness of segmentation models to corruptions in biological imaging

Kesenci, Y.; Le Folgoc, L.; Angelini, E.

2026-08-25 bioinformatics
10.64898/2026.08.21.746302 bioRxiv
Show abstract

Deep-learning-based segmentation algorithms have gained considerable accuracy for processing biological images. In particular, the introduction of large foundation models, novel architectures, and semantically varied datasets now allows for deployment of state-of-the-art models for clean image cohorts with limited re-training or, in the best of cases, in an out-of-the-box fashion. Biological imaging, however, is liable to corruptions that can hinder their deployment. While some methods document their robustness to the most common corruptions, a systematic robustness analysis of the state of the art to the expansive gamut of corruptions in biological imaging remains to be done. We perform this benchmarking by simulating 36 corruption types with varying degradation severity on images sampled from 30 different datasets. Our benchmark accounts both for the variety in biological images and the nature of corruptions. Among other things, our study reveals that performance on clean images does not correlate with overall robustness to image corruptions. In fact, we find that a decade-old method, StarDist, is more robust than many of its more recent foundation-model-based counterparts. We also show in a dedicated representation analysis that the performance of segmentation models collapses in the early layers of the encoding phase.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.