TabVI: Leveraging Lightweight Transformer Architectures to Learn Biologically Meaningful Cellular Representations
Chandrashekar, A.; Gala, R.; tjaernberg, A.; Khullar, S.; Huynh, G.; Gabitto, M. I.
Show abstract
Transformer-based foundation models are changing the landscape of natural language processing (NLP), computer vision, and audio, achieving human-level performance across a variety of tasks. Extending these models to single-cell genomics holds significant potential for revealing the cellular and molecular perturbations associated with disease. However, unlike the sequential structure of language, the functional organization of genes is hierarchical and modular. This fundamental difference necessitates the development of meaningful feature selection strategies to adapt NLP transformer architectures effectively. In contrast to many large-scale foundation models, probabilistic models have shown success in learning complex cellular representations from single-cell datasets. In this work, we present TabVI, a probabilistic deep generative model that leverages tabular transformer architectures to improve latent embedding learning. We validate TabVIs performance in cell type annotation and integration benchmarks. We demonstrate that TabVI improves performance across down-stream tasks and is robust to scaling dataset sizes, producing interpretable, sample-specific feature attention masks. TabVI is a lightweight, scientifically-meaningful, transformer architecture for single-cell analysis that excels where large scale foundation models are less effective.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Single-Cell Multi-Modal GAN (scMMGAN) reveals spatial patterns in single-cell data from triple negative breast cancer 96%
- Hierarchical confounder discovery in the experiment-machine learning cycle 95%
- Generating hard-to-obtain information from easy-to-obtain information: applications in drug discovery and clinical inference 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.