AdaGeneBudget: Cell-Adaptive Gene-Token Allocation for Efficient Single-Cell Foundation Models
Kim, D.; Hwang, U.
Show abstract
Single-cell foundation models (scFMs) represent each cell using sequences of gene-associated tokens, making embedding extraction increasingly costly as the number of cells and expressed genes grows. Existing input policies typically rely on fixed input budgets, with retained genes determined by random subsampling, model-native ranking, or a fixed dataset-level highly variable gene (HVG) panel. However, they do not jointly determine, for each cell, which genes to retain and how many tokens to allocate. We introduce AdaGeneBudget, a training-free gene-token selection method that combines each genes expression with reference-derived inverse detection frequency and retains the shortest ranked prefix that captures a target fraction of the cells expression-specificity score mass. The resulting cell-specific budget is bounded by predefined minimum and maximum lengths, requires no cell-type labels, and leaves the pretrained backbone unchanged. We evaluated AdaGeneBudget in a frozen-backbone inference setting using pretrained scGPT and Geneformer models on Kang and PBMC reference-mapping tasks, with an additional scPRINT comparison against its official HVG policy and an expressed-only HVG control. Across four scGPT and Geneformer backbone-dataset pairs, AdaGeneBudget substantially reduced mean gene-token counts and peak GPU memory while increasing embedding-extraction throughput by up to 4.63x. Despite this compression, it preserved native-level aggregate annotation utility and consistently outperformed token-matched random selection. AdaGeneBudget also preserved fine-grained and low-support cell identities and retained lineage-marker programs and stimulation-associated pathway genes under compression. In scPRINT, both HVG controls achieved higher annotation macro-F1, whereas AdaGeneBudget more faithfully preserved the stimulation-induced embedding direction. These results establish biologically informed, cell-adaptive gene-token allocation as a practical complement to architectural and systems-level efficiency methods for applying existing scFMs to new datasets. They also suggest a cell-adaptive input-allocation principle for future models operating under finite token budgets.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- CellFM: a large-scale foundation model pre-trained on transcriptomics of 100 million human cells 95%
- scDREAMER: atlas-level integration of single-cell datasets using deep generative model paired with adversarial classifier 95%
- RegFormer: A Single-Cell Foundation Model Powered by Gene Regulatory Hierarchies 95%
Similar papers in this journal
- Learning multi-cellular representations of single-cell transcriptomics data enables characterization of patient-level disease states 94%
- Identifying maximally informative signal-aware representations of single-cell data using the Information Bottleneck 94%
- scCausalVI disentangles single-cell perturbation responses with causality-aware generative model 93%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.