Impact of Data Quality on Deep Learning Prediction of Spatial Transcriptomics from Histology Images
Hallinan, C.; Lucas, C.-H. G.; Fan, J.
Show abstract
Spatial transcriptomic technologies enable high-throughput quantification of gene expression at specific locations across tissue sections, facilitating insights into the spatial organization of biological processes. However, high costs associated with these technologies have motivated the development of deep learning methods to predict spatial gene expression from inexpensive hematoxylin and eosin-stained histology images. While most efforts have focused on modifying model architectures to boost predictive performance, the influence of training data quality remains largely unexplored. Here, we investigate how variation in molecular and image data quality stemming from differences in spatial transcriptomic technologies impact deep learning-based gene expression prediction from histology images. To identify the aspects of data quality that impact predictive performance, we conducted in silico ablation experiments, which showed that increased sparsity and noise in molecular data degraded predictive performance, while in silico rescue experiments via imputation provided only limited improvements that failed to generalize beyond the test set. Likewise, reduced image resolution can degrade predictive performance and further impacts model interpretability. We further demonstrate that these data quality-driven effects are reproducible across multiple spatial transcriptomics datasets and remain consistent when using alternative feature extractors and model architectures. Overall, our results show how improving data quality provides an orthogonal strategy to tuning model architecture in spatial transcriptomics-based predictive modeling, highlighting the need to account for technology-specific limitations that directly impact data quality when developing predictive methodologies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SHEST: Single-cell-level artificial intelligence from haematoxylin and eosin morphology for cell type prediction and spatial transcriptomics reconstruction 97%
- SpaTM: Topic Models for Inferring Spatially Informed Transcriptional Programs 96%
- BayeSMART: Bayesian Clustering of Multi-sample Spatially Resolved Transcriptomics Data 96%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.