BatchSVG: identifying batch-biased genes in the application of spatially variable gene detection
Shah, K. H.; Hou, C.; Thompson, J. R.; Hicks, S. C.
Show abstract
1A standard task in the analysis of spatially resolved transcriptomics data is to identify spatially variable genes (SVGs). This is most commonly done within one tissue section at a time because the spatial relationships between the tissue sections are typically unknown. However, large-scale spatial atlases are being generated, for example across hundreds of donors, where the goal is to identify a common set of SVGs to use for downstream analyses. One challenge is how to identify and remove SVGs that are associated with a known bias or technical artifact, such as the slide or capture area, which can lead to poor performance in downstream analyses, such as spatial domain detection. Here, we introduce BatchSVG, a tool to identify batch-biased genes in the application of SVG detection. Our approach compares the rank of per-gene deviance under a binomial model (i) with and (ii) without including a covariate in the model that is associated with the known bias or technical artifact. If the rank of a gene changes significantly between these, then we infer that this gene is likely associated with the bias or technical artifact and should be removed from the downstream analysis. We consider two SRT datasets and show how our model can improve the results of downstream analysis.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Integrating gene expression and imaging data across Visium capture areas with visiumStitched 96%
- MitoDelta: identifying mitochondrial DNA deletions at cell-type resolution from single-cell RNA sequencing data 93%
- Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.