Back

Nicheformer: a foundation model for single-cell and spatial omics

Schaar, A. C.; Tejada-Lapuerta, A.; Palla, G.; Gutgesell, R.; Halle, L.; Minaeva, M.; Vornholz, L.; Dony, L.; Drummer, F.; Bahrami, M.; Theis, F. J.

2024-10-26 bioinformatics
10.1101/2024.04.15.589472 bioRxiv
Show abstract

Tissue makeup relies fundamentally on the cellular microenvironment. Spatial single-cell genomics allows probing the underlying cellular interactions in an unbiased, scalable fashion. To learn a unified cell representation that accounts for local dependencies in the cellular microenvironment, we propose Nicheformer, a transformer-based foundation model that combines human and mouse dissociated single-cell and targeted spatial transcriptomics data. Pretrained on over 57 million dissociated and 53 million spatially resolved cells across 73 tissues on cellular reconstruction, the model is fine-tuned on spatial tasks for spatial omics data to decode spatially resolved cellular information. Nicheformer excels in linear-probing and fine-tuning scenarios for a novel set of downstream tasks, in particular spatial composition prediction and spatial label prediction. We further show that existing foundation models trained on dissociated single-cell data alone are not capable of recapitulating the spatial complexity of cells in their microenvironments, indicating that multiscale models are required to understand complex local dependencies at scale. Nicheformer enables the prediction of the spatial context of dissociated cells, allowing the transfer of rich spatial information to scRNA-seq datasets. Overall, Nicheformer sets the stage for the next generation of machine-learning models in spatial single-cell analysis. Extended AbstractTissue makeup and the corresponding orchestration of vital biological activities, ranging from development and differentiation to immune response and regeneration, rely fundamentally on the cellular microenvironment and the interactions between cells. Spatial single-cell genomics allows probing such interactions in an unbiased and, increasingly, scalable fashion. To learn a unified cell representation that accounts for local dependencies in the cellular microenvironment and the underlying cell interactions, we propose to generalize recent foundation modeling approaches for disassociated single-cell transcriptomics to the spatial omics setting. Our model, Nicheformer, is a transformer-based foundation model that combines human and mouse dissociated single-cell and targeted spatial transcriptomics data to learn a cellular representation useful for a large variety of downstream tasks. Nicheformer is pretrained on over 57 million dissociated and 53 million spatially resolved cells across 73 tissues from both human and mouse. Subsequently, the model is fine-tuned on spatial tasks for spatial omics data to decode spatially resolved cellular information. We demonstrate the usefulness of Nicheformer in both linear-probing as well as fine-tuning scenarios on a novel set of spatially-relevant downstream tasks such as spatial density prediction or niche and region label prediction. In particular, we show that Nicheformer enables the prediction of the spatial context of dissociated cells, allowing the transfer of rich spatial information to scRNA-seq datasets. We define a series of novel spatial prediction problems and observe consistent top performance of Nicheformer, demonstrating the advantage of the improved model capacity of the underlying transformer. Additionally, we benchmarked Nicheformer in these tasks against scGPT1, Geneformer2, scVI3 and PCA and show that the Nicheformer architecture excels in these tasks. Altogether, our large-scale resource of more than 110 million cells in a partial spatial context, together with the set of novel spatial learning tasks and the Nicheformer model itself, will pave the way for the next generation of machine-learning models for spatial single-cell analysis.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.