Annotation-efficient classification combining active learning, pre-training and semi-supervised learning for biomedical images
Shetab Boushehri, S.; Bin Qasim, A.; Waibel, D.; Schmich, F.; Marr, C.
Show abstract
Deep learning based classification of biomedical images requires manual annotation by experts, which is time-consuming and expensive. Incomplete-supervision approaches including active learning, pre-training and semi-supervised learning address this issue and aim to increase classification performance with a limited number of annotated images. Up to now, these approaches have been mostly benchmarked on natural image datasets, where image complexity and class balance typically differ considerably from biomedical classification tasks. In addition, it is not clear how to combine them to improve classification performance on biomedical image data. We thus performed an extensive grid search combining seven active learning algorithms, three pre-training methods and two training strategies as well as respective baselines (random sampling, random initialization, and supervised learning). For four biomedical datasets, we started training with 1% of labeled data and increased it by 5% iteratively, using 4-fold cross-validation in each cycle. We found that the contribution of pre-training and semi-supervised learning can reach up to 25% macro F1-score in each cycle. In contrast, the state-of-the-art active learning algorithms contribute less than 5% to macro F1-score in each cycle. Based on performance, implementation ease and computation requirements, we recommend the combination of BADGE active learning, ImageNet-weights pre-training, and pseudo-labeling as training strategy, which reached over 90% of fully supervised results with only 25% of annotated data for three out of four datasets. We believe that our study is an important step towards annotation and resource efficient model training for biomedical classification challenges.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SN-FPN: Self-attention Nested Feature Pyramid Network for Digital Pathology Image Segmentation 96%
- Cell segmentation without annotation by unsupervised domain adaptation based on cooperative self-learning 95%
- The tempest in a cubic millimeter: Image-based refinements necessitate the reconstruction of 3D microvasculature from a large series of damaged alternately-stained histological sections 94%
Similar papers in this journal
- Enhancing Breast Ultrasound Segmentation through Fine-tuning and Optimization Techniques: Sharp Attention UNet 96%
- Semantic Segmentation of HeLa Cells: An Objective Comparison between one Traditional Algorithm and Three Deep-Learning Architectures 95%
- DMENet: Diabetic Macular Edema Diagnosis using Hierarchical Ensemble of CNNs 95%
Similar papers in this journal
- A novel interpretable deep transfer learning combining diverse learnable parameters for improved T2D prediction based on single-cell gene regulatory networks 96%
- Probabilistic Brain MR Image Transformation Using Generative Models 95%
- ARA: accurate, reliable and active histopathological image classification framework with Bayesian deep learning 95%
Similar papers in this journal
- BenchXAI: Comprehensive Benchmarking of Post-hoc Explainable AI Methods on Multi-Modal Biomedical Data 97%
- MultiHeadGAN: A Deep Learning Method for Low Contrast Retinal Pigment Epithelium Cells Segmentation in Fluorescent Flatmount Microscopy Images 95%
- Deep Multimodal Graph-Based Network for Survival Prediction from Highly Multiplexed Images and Patient Variables 94%
Similar papers in this journal
- STAMP: Simultaneous Training and Model Pruning for Low Data Regimes in Medical Image Segmentation 97%
- A Framework for Falsifiable Explanations of Machine Learning Models with an Application in Computational Pathology 94%
- Transformer with Convolution and Graph-Node co-embedding: An accurate and interpretable vision backbone for predicting gene expressions from local histopathological image 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.