Mitigating Bias in Spatial Transcriptomic Pipelines via Human Feedback
Boyeau, P.; Bates, S.; Ergen, C.; Jordan, M. I.; Yosef, N.
Show abstract
Biological discovery from experimental data, particularly large-scale assays, requires extensive preprocessing, during which raw outputs (e.g., images, sequences) are processed into structured forms that are more amenable to analysis. While statistical methods for such processed data are at the core of computational biology, the problem of coping with uncertainties introduced during preprocessing is a significant and underexplored issue. We address this issue in the context of differential expression analysis in spatial transcriptomics, which depends on a series of preprocessing steps, including demarcation of cell regions (segmentation), quantification of gene expression in cells, and cell-type annotation. We introduce Corrected Spatial Differential Expression (CSDE), a method that builds on Prediction-Powered Inference to leverage a small set of expert-validated data points (cells) to account for uncertainty due to preprocessing errors. Using two case studies, we demonstrate that CSDE produces more reliable and calibrated estimates of differential expression compared to the prevalent approach that neglects the impact of preprocessing. CSDE incorporates an efficient workflow to generate the required expert-annotated data, and is available as open-source at https://github.com/YosefLab/CSDE.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Automated assignment of cell identity from single-cell multiplexed imaging and proteomic data 98%
- Belayer: Modeling discrete and continuous spatial variation in gene expression from spatially resolved transcriptomics 96%
- Uncovering the spatial landscape of molecular interactions within the tumor microenvironment through latent spaces 96%
Similar papers in this journal
- Randomized Spatial PCA (RASP): a computationally efficient method for dimensionality reduction of high-resolution spatial transcriptomics data 97%
- A Bayesian method to infer copy number clones from single-cell RNA and ATAC sequencing 96%
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.