How many specimens make a sufficient training set for automated 3D feature extraction?
Mulqueeney, J. M.; Searle-Barnes, A.; Brombacher, A.; Sweeney, M.; Goswami, A.; Ezard, T.
Show abstract
Deep learning has emerged as a robust tool for automating feature extraction from 3D images, offering an efficient alternative to labour-intensive and potentially biased manual image segmentation methods. However, there has been limited exploration into the optimal training set sizes, including assessing whether artificial expansion by data augmentation can achieve consistent results in less time and how consistent these benefits are across different types of traits. In this study, we manually segmented 50 planktonic foraminifera specimens from the genus Menardella to determine the minimum number of training images required to produce accurate volumetric and shape data from internal and external structures. The results reveal unsurprisingly that deep learning models improve with a larger number of training images and that data augmentation can enhance network accuracy by up to 8.0%. Notably, predicting both volumetric and shape measurements for the internal structure poses a greater challenge compared to the external structure, due to low contrast between different materials and increased geometric complexity. These results provide novel insight into optimal training set sizes for precise image segmentation of diverse traits and highlight the potential of data augmentation for enhancing multivariate feature extraction from 3D images. Subject CategoryLife Sciences - Computer Science Subject Areascomputational biology, Artificial Intelligence
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Fluorescence Microscopy Datasets for Training Deep Neural Networks 93%
- Artifact-free whole-slide imaging with structured illumination microscopy and Bayesian image reconstruction 92%
- CellBinDB: A Large-Scale Multimodal Annotated Dataset for Cell Segmentation with Benchmarking of Universal Models 92%
Similar papers in this journal
- A new tool for 3D segmentation of computed tomography data: Drishti Paint and its applications 94%
- New Interactive Machine Learning Tool for Marine Image Analysis 94%
- A New Straightforward Method for Automated Segmentation of Trabecular Bone from Cortical Bone in Diverse and Challenging Morphologies 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.