Back

Zero-Shot, Big-Shot, Active-Shot - How to estimate cell confluence, lazily

Joas, M. J.; Freund, D.; Haase, R.; Rahm, E.; Ewald, J.

2025-01-21 bioinformatics
10.1101/2025.01.17.633501 bioRxiv
Show abstract

Mesenchymal stem cell therapy shows promising results for difficult-to-treat diseases, but standardized manufacturing requires robust quality control through automated cell confluence monitoring. While deep learning can automate confluence estimation, research on cost-effective dataset curation and the role of foundation models in this task remains limited. We systematically investigate the most effective strategies for confluence estimation, focusing on active learning-based dataset curation, goal-specific labeling, and leveraging foundation models for zero-shot inference. Here, we show that zero-shot inference with the Segment Anything Model (SAM) achieves excellent confluence estimation without any task-specific training, outperforming fine-tuned smaller models. Further, our findings demonstrate that active learning does not significantly improve model dataset curation compared to random selection in homogeneous cell datasets. We show that goal-specific, simplified labeling strategies perform comparably to precise annotations while substantially reducing annotation effort. These results challenge common assumptions about dataset curation: neither active learning nor extensive fine-tuning provided significant benefits for our specific use case. Instead, we found that leveraging SAMs zero-shot capabilities and targeted labeling strategies offers the most cost-effective approach to automated confluence estimation. Our work provides practical guidelines for implementing automated cell monitoring in MSC manufacturing, demonstrating that extensive dataset curation may be unnecessary when foundation models can effectively handle the task out of the box.

Matching journals

The top 9 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.