Back

SynSeg: Generating Synthetic Datasets for Accurate Subcellular Segmentation with U-net

Guo, Z.; XU, K.; Wang, Z.; Ke, J.; Huang, J.; Ye, Y.; Ou, G.

2025-02-08 bioinformatics
10.1101/2025.02.07.637194 bioRxiv
Show abstract

Accurate segmentation of subcellular components is crucial for understanding cellular processes, but traditional methods struggle with noise and complex structures. Convolutional neural networks improve accuracy but require large, time-consuming, and biased manually annotated datasets. Here, we developed SynSeg, a pipeline that generates synthetic training data to train a U-net model for subcellular structure segmentation, eliminating the need for manual annotation. SynSeg leverages synthetic datasets with variations in intensity, morphology, and signal distribution to deliver context-aware segmentations, even in challenging imaging conditions. We demonstrate SynSegs superior performance in segmenting vesicles and cytoskeletal filaments from culture cells and live C. elegans, outperforming traditional methods such as Otsus thresholding, ILEE, and FilamentSensor 2.0. Additionally, SynSeg effectively quantified disease-associated microtubule morphology in live cells, uncovering structural defects caused by mutant Tau proteins linked to neurodegenerative diseases. These results highlight the potential of synthetic data-driven approaches to advance biological segmentation and enhance microscopy techniques. Significance StatementThis study introduces a novel approach for accurately segmenting cellular structures, such as microtubules and vesicles, using synthetic datasets and advanced deep learning techniques. By leveraging a U-Net model trained on thousands of artificially generated images, our method eliminates the need for labor-intensive experimental data and simplifies the data creation process. Importantly, it incorporates noise and variability into the training datasets to make the model more robust and biologically relevant. Our findings demonstrate that the model can successfully identify cellular components, paving the way for its application in real-world microscopy images. This innovation has the potential to accelerate discoveries in cell biology by providing an efficient, scalable tool for analyzing complex cellular structures, even in challenging imaging conditions.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.