Back

Extracting deep learning based morphology segmentation footprint for boar sperm cells

Park, J.; Ratka, M.; Biswas, A.; Shofner, I.; Kerns, K.; Sarkar, A.

2026-08-13 bioinformatics
10.64898/2026.08.07.743571 bioRxiv
Show abstract

Reliable delineation of the head and tail of swine spermatozoa supports automated assessment of boar semen quality, from morphometric measurement to the quality control of insemination doses. In practice this relies on fluorescent staining, which adds chemistry, cost, and delay to every acquisition and labels only the nucleus. Recent work coupling imaging flow cytometry with machine learning has advanced rapidly, yet the segmentation stage still depends on a stained channel at inference and resolves the head alone. We present a supervised encoder decoder network that segments boar spermatozoa from brightfield images acquired on an Amnis ImageStream Mark II with no stain at inference. Training labels derive from the Hoechst 33342 nuclear channel (Ch7), recorded in registration with brightfield (Ch1); the dye serves only as an annotation source, and the network sees Ch1 alone. The best semantic segmentation model reaches a Dice coefficient of 0.940 on held-out cells. For comparison we evaluate a classical morphological pipeline, four further semantic segmentation models spanning three decoder families and two ImageNet-pretrained backbones, and two zero-shot pipelines built on the Segment Anything Model 2 (SAM 2), prompted either by a dilated box around the predicted head mask or by head and tail boxes emitted by a Gemma 4 Vision Language Model (VLM). The zero-shot route scores 0.637 against Ch7 but labels the tail, which the fluorescence protocol cannot. Cells scoring worst under the supervised model proved to be mostly registration failures rather than segmentation failures, as Ch7 is displaced relative to Ch1. Manual screening for this drift is infeasible at dataset scale, so we propose a flagging system that marks any Dice below 0.792, two standard deviations below the mean, and pairs it with a zero-shot pipeline in which a VLM l and SAM 2 cross-check the flagged cell before human review.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.