BEGA-UNet: Boundary-Explicit Guided Attention U-Net with Multi-Scale Feature Aggregation for Colonoscopic Polyp Segmentation
Tong, T.; Zhang, W.; Zu, W.
Show abstract
Accurate polyp segmentation from colonoscopy images is critical for colorectal cancer prevention, yet the generalization of deep learning models under domain shift remains insufficiently explored. We propose Boundary-Explicit Guided Attention U-Net (BEGA-UNet), a boundary-aware segmentation architecture that introduces explicit edge modeling as a structural inductive bias to enhance both segmentation accuracy and cross-domain robustness. The framework integrates three components: an Edge-Guided Module (EGM) with learnable Sobel-initialized operators to capture boundary cues, a Dual-Path Attention (DPA) module that processes channel and spatial attention in parallel, and a Multi-Scale Feature Aggregation (MSFA) module to encode contextual information across multiple receptive fields. Evaluated on the combined Kvasir-SEG and CVC-ClinicDB benchmarks, BEGA-UNet achieves 88.53% Dice and 82.51% IoU, outperforming representative convolutional and transformer-based baselines. More importantly, cross-dataset evaluation demonstrates strong robustness under domain shift, with BEGA-UNet retaining 83.2% of its in-distribution performance--substantially higher than U-Net (64.5%), Attention U-Net (47.5%), and TransUNet (53.1%). In a zero-shot setting on an entirely unseen dataset, the model further maintains 72.6% performance retention. Comprehensive ablation studies indicate that explicit boundary modeling plays a central role in improving generalization, while multi-scale context aggregation further stabilizes performance across domains. Feature distribution analyses support this observation by showing that edge-oriented representations exhibit markedly reduced cross-domain variability compared to appearance-driven features. Overall, BEGA-UNet provides an effective and interpretable solution for robust polyp segmentation, demonstrating that explicit boundary modeling serves as a critical inductive bias for ensuring reliability under clinical domain shifts.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Identifying perturbations that boost T-cell infiltration into tumours via counterfactual learning of their spatial proteomic profiles 92%
- Is the brain macroscopically linear? A system identification of resting state dynamics 90%
- Aggregation Tool for Genomic Concepts (ATGC): A deep learning framework for somatic mutations and other sparse genomic measures 88%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Clinical Validation of Saliency Maps for Understanding Deep Neural Networks in Ophthalmology 95%
- A Framework for Falsifiable Explanations of Machine Learning Models with an Application in Computational Pathology 93%
- Spatial Transcriptomics Expression Prediction from Histopathology Based on Cross-Modal Mask Reconstruction and Contrastive Learning 93%
Similar papers in this journal
- Dual Adversarial Deconfounding Autoencoder for joint batch-effects removal from multi-center and multi-scanner radiomics data 94%
- Aggregation of Cohorts for Histopathological Diagnosis with Deep Morphological Analysis 94%
- DAMM for the detection and tracking of multiple animals within complex social and environmental settings 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.