How different AI models understand cells differently
Zhao, Y.; Sun, D.; Hao, M.; Xiong, Y.; Li, C.; Gong, T.; Wei, L.; Zhang, X.
Show abstract
AI single-cell foundation models (scFMs) are believed to be able to learn essential relations in cell transcriptomics with the attention modules in Transformer, but there is no method to reveal what they actually learned. We observed that different models may grasp different aspects of relations. To unravel the mystery, we propose scGeneLens, a framework for dissecting how scFMs perceive cells. We employed a sparse block attention to replace the original attention mechanism to concentrate attentions into a few dominant gene-gene relations, used attention propagation to trace how the relations propagate across Transformer layers, and used integrated gradients to disentangle the relative contributions of gene identity and expression in cell representations. We applied it to scFoundation and scGPT and show that they exhibit pluralistic perceptions of cells: scFoundation emphasizes relations among cell-type marker genes, resulting in stronger cell-type separability, whereas scGPT focuses more on genes involved in shared cellular pathways and core biological activities, leading to representations that generalize across conditions. The framework provides a unified lens for probing what scFMs learn about cells and offers actionable insights for the design of future cellular foundation models. Our code can be seen in https://anonymous.4open.science/r/scGeneLens-B771/.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Hypergraph factorisation for multi-tissue gene expression imputation 96%
- Interpreting Neural Networks for Biological Sequences by Learning Stochastic Masks 95%
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.