Back

Reference-based cell type matching of spatial transcriptomics data

Zhang, Y.; Miller, J. A.; Park, J.; Lelieveldt, B. P.; Long, B.; Abdelaal, T.; Aevermann, B. D.; Biancalani, T.; Comiter, C.; Dzyubachyk, O.; Mattsson Langseth, C.; Petukhov, V.; Scalia, G.; Vaishnav, E. D.; Zhao, Y.; Lein, E. S.; Scheuermann, R. H.

2022-05-12 bioinformatics
10.1101/2022.03.28.486139 bioRxiv
Show abstract

With the advent of multiplex fluorescence in situ hybridization (FISH) and in situ RNA sequencing technologies, spatial transcriptomics analysis is advancing rapidly. Spatial transcriptomics provides spatial location and pattern information about cells in tissue sections at single cell resolution. Cell type classification of spatially-resolved cells can also be inferred by matching the spatial transcriptomics data to reference single cell RNA-sequencing (scRNA-seq) data with cell types determined by their gene expression profiles. However, robust cell type matching of the spatial cells is challenging due to the intrinsic differences in resolution between the spatial and scRNA-seq data. In this study, we systematically evaluated six computational algorithms for cell type matching across four spatial transcriptomics experimental protocols (MERFISH, smFISH, BaristaSeq, and ExSeq) conducted on the same mouse primary visual cortex (VISp) brain region. We find that while matching results of individual algorithms vary to some degree, they also show agreement to some extent. We present two ensembl meta-analysis strategies to combine the individual matching results and share the consensus matching results in the Cytosplore Viewer (https://viewer.cytosplore.org) for interactive visualization and data exploration. The consensus matching can also guide spot-based spatial data analysis using SSAM, allowing segmentation-free cell type assignment.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.