Sparse Autoencoders Reveal Interpretable Features in Single-Cell Foundation Models
Pedrocchi, F.; Barkmann, F.; Joudaki, A.; Boeva, V.
Show abstract
Single-cell foundation models (scFMs) hold promise for applications in cell type annotation, data integration, and prediction of the effects of cell perturbations, but their internal mechanisms remain poorly understood. We investigate the structure of these models by training sparse autoencoders (SAEs) on the hidden representations of three widely used scFMs: scGPT, scFoundation, and Geneformer. The learned features reveal diverse and complex biological and technical signals, which emerge even in pre-trained models. We also observe that the encoding of this information differs between scFMs with distinct training protocols and architectures. Finally, we demonstrate that SAE-derived features are functionally related to model behavior and can be intervened upon to reduce unwanted technical effects while steering model outputs to preserve the core biological signal. These findings provide a path toward more interpretable and controllable singlecell foundation models.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deep learning of gene interactions from single cell time-course expression data 95%
- scValue: value-based subsampling of large-scale single-cell transcriptomic data for machine and deep learning tasks 95%
- Evaluation of out-of-distribution detection methods for data shifts in single-cell transcriptomics 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.