CellPatch: a Highly Efficient Foundation Model for Single-Cell Transcriptomics with Heuristic Patching
Wu, H.-J.; Zheng, X.; Ma, Z.; Zhu, H.; Yuan, Y.; Yang, J.; Cai, K.; Wei, N.; Zhang, S.; Wang, L.; Wenjie, J.; Sun, Y.; Wang, Y.-J.; Liu, A.; Lai, F.
Show abstract
The rapid advancement of foundation models has significantly enhanced the analysis of single-cell omics data, enabling researchers to gain deeper insights into the complex interactions between cells and genes across diverse tissues. However, existing foundation models often exhibit excessive complexity, hindering their practical utility for downstream tasks. Here, we present CellPatch, a lightweight foundation model that leverages the strengths of the cross-attention mechanism and patch tokenization to reduce model complexity while extracting efficient biological representations. Comprehensive evaluations conducted on single-cell RNA-sequencing datasets across multiple organs and tissue states demonstrate that CellPatch achieves state-of-the-art performance in downstream analytical tasks while maintaining ultra-low computational costs during both pretraining and finetuning phases. Moreover, the flexibility and scalability of CellPatch allow it to serve as a general framework that can be incorporated with other well established single-cell analysis software, thereby enhancing their performance through transfer learning on diverse downstream tasks.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- INSCT: Integrating millions of single cells using batch-aware triplet neural networks 97%
- Inferring spatial single-cell-level interactions through interpreting cell state and niche correlations learned by self-supervised graph transformer 97%
- Simultaneous dimensionality reduction and integration for single-cell ATAC-seq data using deep learning 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.