Principled PCA separates signal from noise in omics count data
Stanley, J. S.; Yang, J.; Li, R.; Lindenbaum, O.; Kobak, D.; Landa, B.; Kluger, Y.
Show abstract
Principal component analysis (PCA) is indispensable for processing high-throughput omics datasets, as it can extract meaningful biological variability while minimizing the influence of noise. However, the suitability of PCA is contingent on appropriate normalization and transformation of count data, and accurate selection of the number of principal components; improper choices can result in the loss of biological information or corruption of the signal due to excessive noise. Typical approaches to these challenges rely on heuristics that lack theoretical foundations. In this work, we present Biwhitened PCA (BiPCA), a theoretically grounded framework for rank estimation and data denoising across a wide range of omics modalities. BiPCA overcomes a fundamental difficulty with handling count noise in omics data by adaptively rescaling the rows and columns - a rigorous procedure that standardizes the noise variances across both dimensions. Through simulations and analysis of over 100 datasets spanning seven omics modalities, we demonstrate that BiPCA reliably recovers the data rank and enhances the biological interpretability of count data. In particular, BiPCA enhances marker gene expression, preserves cell neighborhoods, and mitigates batch effects. Our results establish BiPCA as a robust and versatile framework for high-throughput count data analysis.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 97%
- scINSIGHT for interpreting single-cell gene expression from biologically heterogeneous data 96%
- CMOT: Cross Modality Optimal Transport for multimodal inference 96%
Similar papers in this journal
- CellScope: High-Performance Cell Atlas Workflow with Tree-Structured Representation 96%
- Probabilistic embedding, clustering, and alignment for integrating spatial transcriptomics data with PRECAST 96%
- scMODAL: A general deep learning framework for comprehensive single-cell multi-omics data alignment with feature links 96%
Similar papers in this journal
- Randomized Spatial PCA (RASP): a computationally efficient method for dimensionality reduction of high-resolution spatial transcriptomics data 97%
- BARcode DEmixing through Non-negative Spatial Regression (BarDensr) 96%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 96%
Similar papers in this journal
- scPrisma: inference, filtering and enhancement of periodic signals in single-cell data using spectral template matching 97%
- Multi-resolution deconvolution of spatial transcriptomics data reveals continuous patterns of inflammation 97%
- scJoint: transfer learning for data integration of atlas-scale single-cell RNA-seq and ATAC-seq 97%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.