Back

Towards Generalizability and Robustness in Biological Object Detection in Electron Microscopy Images

GIANNIOS, K.; Chaurasia, A.; Bueno, C.; Riesterer, J. L.; Pagano, L.; Lo, T. P.; Thibault, G.; Gray, J. W.; Song, X.; DeLaRosa, B.

2023-11-27 bioengineering
10.1101/2023.11.27.568889 bioRxiv
Show abstract

Machine learning approaches have the potential for meaningful impact in the biomedical field. However, there are often challenges unique to biomedical data that prohibits the adoption of these innovations. For example, limited data, data volatility, and data shifts all compromise model robustness and generalizability. Without proper tuning and data management, deploying machine learning models in the presence of unaccounted for corruptions leads to reduced or misleading performance. This study explores techniques to enhance model generalizability through iterative adjustments. Specifically, we investigate a detection tasks using electron microscopy images and compare models trained with different normalization and augmentation techniques. We found that models trained with Group Normalization or texture data augmentation outperform other normalization techniques and classical data augmentation, enabling them to learn more generalized features. These improvements persist even when models are trained and tested on disjoint datasets acquired through diverse data acquisition protocols. Results hold true for transformerand convolution-based detection architectures. The experiments show an impressive 29% boost in average precision, indicating significant enhancements in the models generalizibality. This underscores the models capacity to effectively adapt to diverse datasets and demonstrates their increased resilience in real-world applications.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.