BaySiCle: A Bayesian Inference joint kNN method for imputation of single-cell RNA-sequencing data making use of local effect
Singh, A. N.
Show abstract
There is a marked technical variability and a high amount of missing observations in the single-cell data that we obtain from experiments. Apart from that clearly each of the batch of experiments have a batch effect on every cell in the batch. This batch effect can be taken into advantage for dealing with imputation, given that all the cells in a given batch belong to the same tissue. Here we introduce BaySiCle, a novel Bayesian inference based method combined with k-nearest neighbors algorithm for the imputation of missing data in scRNA-seq counts. The priors are found out based on expression value across cells for all the single cells of the same batch. We demonstrate using sample scRNA-seq datasets and simulated expression data that BaySiCle allows robust imputation of missing values generating realistic transcript distributions that match single molecule fluorescence in situ hybridization measurements. By using priors as obtained by the dataset structures in the not just the experimental set-up batch, but also the same group of cells, BaySiCle improves accuracy of imputation to be that much closer to its similar alternatives. Availability and implementationThe Python Jupyter notebook BaySiCle is published on GitHub GitHub - abinarain/BaySiCel: Single Cell Data Imputation using Bayesian statistics and kNN
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Deepprune: Learning efficient and interpretable convolutional networks through weight pruning for predicting DNA-protein binding 93%
- SSAM-lite: a light-weight web app for rapid analysis of spatially resolved transcriptomics data 93%
- Accelerated matrix-vector multiplications for matrices involving genotype covariates with applications in genomic prediction 93%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.