Back

SpaFun: Discovering Domain-specific Spatial Expression Patterns and New Disease-Relevant Genes using Functional Principal Component Analysis

Jiang, X.; Guo, Y.; Guo, L.; Zhong, L.; Wang, J.; Xiao, G.; Li, Q.; Xu, L.

2025-02-21 genomics
10.1101/2025.02.17.638766 bioRxiv
Show abstract

SpaFun is a novel, non-model-based method developed to address limitations in existing spatially variable gene (SVG) detection techniques, particularly for large-scale spatially resolved transcriptomics (SRT) datasets. These limitations include computational inefficiency, limited statistical power with increasing data size, and the inability to capture spatial heterogeneity and co-expression patterns among genes. Built on functional principal component analysis (fPCA), SpaFun identifies domain-representative genes (DRGs) with significantly better computational efficiency and greater statistical power while accounting for spatial heterogeneity and co-expression patterns among genes. We applied SpaFun to three SRT datasets and demonstrated that SpaFun outperformed state-of-the-art algorithms for identifying representative genes for tumor regions (e.g., DESeq, edgeR, and limma), as well as recently developed novel algorithms designed for spatial omics to identify the representative genes (e.g., SPARK and CSIDE). This highlights SpaFuns ability to accurately identify genes most representative of each spatial domain (e.g., tumor, immune, or stroma regions). By uncovering novel disease-relevant genes overlooked by existing algorithms, SpaFun could provide insights into new molecular mechanisms and propose innovative therapeutic strategies to improve patient outcomes. Key PointsO_LISpaFun is the first method dedicated to identifying DRGs, capturing spatially representative expression patterns within annotated tissue regions, setting it apart from all traditional SVG detection methods. C_LIO_LIIt leverages fPCA to model gene expression as a function of spatial location, avoiding reliance on predefined spatial distribution assumptions. C_LIO_LIThe non-model-based framework ensures compatibility with different SRT platforms and experimental designs, making it a scalable and widely applicable tool for SRT research. C_LI

Published in Briefings in Bioinformatics (predicted rank #2) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.