Back

CLM-X: A multimodal single-cell foundation model with flexible multi-way Transformer for unified scRNA-seq and scATAC-seq analysis

Li, B.; Liu, Z.; Wang, Z.; Xu, Z.; Li, Y.; Sha, C.; Li, X.

2026-02-18 genomics
10.64898/2026.02.17.704943 bioRxiv
Show abstract

AO_SCPLOWBSTRACTC_SCPLOWAdvances in single-cell multimodal profiling have enabled a more systematic analysis of cellular biology, yet the rapid accumulation of large-scale, heterogeneous datasets poses substantial challenges for integrative analysis. Recently, Transformer-based cell language models (CLMs) are becoming powerful foundational tools for learning transferable cell representations from unimodal single-cell datasets. However, a unified and flexible multimodal foundation model for joint modeling of scRNA-seq and scATAC-seq datasets remains underexplored. Here, we present CLM-X, a multimodal single-cell foundation model built on multiway Transformer architecture. CLM-X employs a harmonized tokenization design together with a stage-wise masked reconstruction pretraining strategy, enabling unified modeling of RNA-only, ATAC-only, and paired RNA-ATAC input within a single Transformer-based framework. We pretrain CLM-X on million-scale unimodal and multimodal datasets, and systematically evaluate its transferability on five downstream tasks including batch correction, modality integration, cross-modal translation, cell type annotation, and perturbation prediction. Across comprehensive benchmarks on 10 datasets, CLM-X consistently outperforms existing multimodal methods and unimodal foundation models, with particularly clear advantages in RNA-ATAC cross-modal translation and genetic-perturbation-response prediction. Overall, CLM-X establishes a unified and flexible multimodal foundation model for integrative analysis of scRNA-seq and scATAC-seq datasets, advancing a more robust, comprehensive, and biological interpretable single-cell analysis beyond current multimodal fusion approaches and unimodal foundation models.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.