Towards a Cytometry Foundation Model: Interpretable Sample-level Predictive Modelling via Pretrained Transformers
Zhuang, Z.; Mashford, B. S.; Zheng, L.; Andrews, T. D.
Show abstract
Foundation models have transformed scientific data modelling across domains, yet flow cytometry has lacked one. Despite the abundance of high-dimensional cellular data, automated analysis remains bottlenecked by marker variability: prior studies are typically confined to fixed marker panels and homogeneous data, limiting scalability and generalisation due to architectural constraints. We present the Generalised Pretrained Cytometry Transformer (GPCT), an interpretable framework designed to learn from heterogeneous marker panels for sample-level predictive modelling. Through a novel cytometry-specific pretraining regime, GPCT learns transferable cellular representations that achieve high classification accuracy across diverse datasets. Notably, pretraining significantly boosts performance on data-scarce downstream tasks, marking a pivotal step towards a cytometry foundation model. Furthermore, GPCT maintains interpretability and identifies the specific cell subsets most influential to its predictions. This enables direct biological validation of learned patterns and provides a data-driven basis for refining traditional gating strategies.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- multiDGD: A versatile deep generative model for multi-omics data 96%
- scSemiProfiler: Advancing Large-scale Single-cell Studiesthrough Semi-profiling with Deep Generative Models andActive Learning 96%
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 96%
Similar papers in this journal
Similar papers in this journal
- An in-depth comparison of linear and non-linear joint embedding methods for bulk and single-cell multi-omics 96%
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 96%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.