Back

SCARF: Single Cell ATAC-seq and RNA-seq Foundation model

Liu, G.; Zhao, Y.; Zhao, Y.; Wang, T.; Cai, Q.; Wang, X.; Wen, Z.; Lin, L.; Yang, G.; Chen, J.

2025-04-13 bioinformatics
10.1101/2025.04.07.647689 bioRxiv
Show abstract

Recent advances in single-cell multi-omics have provided unprecedented insights into gene regulation by jointly profiling transcriptomic (scRNA-seq) and chromatin accessibility (scATAC-seq) landscapes. However, the inherent heterogeneity and high dimensionality of these multimodal data present significant challenges for effective integration and downstream analysis. Foundation models have demonstrated strong representation learning capabilities for scRNA-seq or scATAC-seq data. So far, however, no model has been specifically developed for the integrative analysis of these two modalities. Here, we introduce SCARF, a single cell ATAC-seq and RNA-seq foundation model. SCARF is pre-trained on X-Omics, the largest curated collection of single-cell multi-omics data to date, comprising over 2.7 million cells across multiple tissues and species. The model utilizes a Mamba architecture for efficiently capturing long-context relationships between genes and between accessible regions. Modality-specific and shared features are learned by the model through self-supervised learning and contrastive learning, respectively. SCARF achieves state-of-the-art performance on multiple downstream tasks, including cell representation, cell matching, and cross-omics translation. Furthermore, SCARF enables few-shot cell type annotation, demonstrating strong generalizability across previously unseen datasets. These results highlight the power of foundation models for advancing integrative analysis of single cell multi-omics data, with broad applications in important tasks including cellular characterization, gene or genomic perturbation analysis, and regulation network analysis.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.