Diffusion on PCA-UMAP manifold captures a well-balance of local, global, and continuum structures to denoise single-cell RNA sequencing data
Padron, C. J.; Resendis-Antonio, O.
Show abstract
Single-cell transcriptomics (scRNA-seq) is becoming a technology that is transforming biological discovery in many fields of medicine. Despite its impact in many areas, scRNASeq is technologically and experimentally limited by the inefficient transcript capture and the high rise of noise sources. For that reason, imputation methods were designed to denoise and recover missing values. Many imputation methods (e.g., neighbor averaging or graph diffusion) rely on k nearest neighbor graph construction derived from a mathematical space as a low-dimensional manifold. Nevertheless, the construction of mathematical spaces could be misleading the representation of densities of the distinct cell phenotypes due to the negative effects of the curse of dimensionality. In this work, we demonstrated that the imputation of data through diffusion approach on PCA space favor over-smoothing when increases the dimension of PCA and the diffusion parameters, such k-NN (k-nearest neighbors) and t (value of the exponentiation of the Markov matrix) parameters. In this case, the diffusion on PCA space distorts the cell neighborhood captured in the Markovian matrix creating an artifact by connecting densities of distinct cell phenotypes, even though these are not related phenotypically. In this situation, over-smoothing of data is due to the fact of shared information among spurious cell neighbors. Therefore, it can not account for more information on the variability (from principal components) or nearest neighbors for a well construction of a cell-neighborhood. To solve above mentioned issues, we propose a new approach called sc-PHENIX( single cell-PHEnotype recovery by Non-linear Imputation of gene eXpression) which uses PCA-UMAP initialization for revealing new insights into the recovered gene expression that are masked by diffusion on PCA space. sc-PHENIX is an open free algorithm whose code and some examples are shown at https://github.com/resendislab/sc-PHENIX.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 96%
- scHiCTools: a computational toolbox for analyzing single cell Hi-C data 96%
- G2S3: a gene graph-based imputation method for single-cell RNA sequencing data 96%
Similar papers in this journal
- Stardust: improving spatial transcriptomics data analysis through space aware modularity optimization based clustering. 95%
- Smash++: an alignment-free and memory-efficient tool to find genomic rearrangements 95%
- Imputing missing RNA-seq data from DNA methylation by using transfer learning based-deep neural network 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.