Learning context-aware, distributed gene representations in spatial transcriptomics with SpaCEX
Sun, X.; Xu, Y.; Li, W.; Huang, M.; Wang, Z.; Chen, J.; Wu, H.
Show abstract
Distributed gene representations are pivotal in data-driven genomic research, offering a structured way to understand the complexities of genomic data and providing foundation for various data analysis tasks. Current gene representation learning methods demand costly pretraining on heterogeneous transcriptomic corpora, making them less approachable and prone to over-generalization. For spatial transcriptomics (ST), there is a plethora of methods for learning spot embeddings but serious lacking method for generating gene embeddings from spatial gene profiles. In response, we present SpaCEX, a pioneer cost-effective self-supervised learning model that generates gene embeddings from ST data through exploiting spatial genomic "context" identified as spatially co-expressed gene groups. SpaCEX-generated gene embeddings (SGE) feature in context-awareness, rich semantics, and robustness to cross-sample technical artifacts. Extensive real data analyses reveal biological relevance of SpaCEX-identified genomic contexts and validate functional and relational semantics of SGEs. We further develop a suite of SGE-based computational methods for a range of key downstream objectives: identifying disease-associated genes and gene-gene interactions, pinpointing genes with designated spatial expression patterns, enhancing transcriptomic coverage of FISH-based ST, detecting spatially variable genes, and improving spatial clustering. Extensive real data results demonstrate these methods superior performance, thereby affirming the potential of SGEs in facilitating various analytical task. Significance StatementSpatial transcriptomics enables the identification of spatial gene relationships within tissues, providing semantically rich genomic "contexts" for understanding functional interconnections among genes. SpaCEX marks the first endeavor to effectively harnesses these contexts to yield biologically relevant distributed gene representations. These representations serve as a powerful tool to greatly facilitate the exploration of the genetic mechanisms behind phenotypes and diseases, as exemplified by their utility in key downstream analytical tasks in biomedical research, including identifying disease-associated genes and gene interactions, in silico expanding the transcriptomic coverage of low-throughput, high-resolution ST technologies, pinpointing diverse spatial gene expression patterns (co-expression, spatially variable pattern, and patterns with specific expression levels across tissue domains), and enhancing tissue domain discovery.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Predicting MammaPrint Recurrence Risk from Breast Cancer Pathological Images Using a Weakly Supervised Transformer 95%
- DeDoc2 identifies and characterizes the hierarchy and dynamics of chromatin TAD-like domains in the single cells 94%
- NanoLoop: A deep learning framework leveraging Nanopore sequencing for chromatin loop prediction 94%
Similar papers in this journal
- stGCL: A versatile cross-modality fusion method based on multi-modal graph contrastive learning for spatial transcriptomics 96%
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 96%
- A Message Passing Framework for Precise Cell State Identification with scClassify2 95%
Similar papers in this journal
Similar papers in this journal
- Probabilistic cell/domain-type assignment of spatial transcriptomics data with SpatialAnno 96%
- Cell type identification in spatial transcriptomics data can be improved by leveraging cell-type-informative paired tissue images using a Bayesian probabilistic model. 96%
- scIGANs: single-cell RNA-seq imputation using generative adversarial networks 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.