Low-Rank Full Matrix Factorization for dropout imputation in single cell RNA-seq and benchmarking with imputation algorithms for downstream applications
Huang, J.; Chow, A. C. M.; Tang, N. L.-s.; Yam, S. C.
Show abstract
BackgroundWhile single cell RNA sequencing becomes a powerful technology, the presence of the large number of zero counts represents a challenge for both wet-lab processing and data analysis. Imputation of these dropouts can now be performed by three categories of algorithms: Model or smoothing, Matrix theory or Deep learning. However, two fundamental questions remain unsettled: (1) whether imputation should be performed; (2) which imputation algorithm to use with various downstream applications. Notably, imputation is not commonly used in real scRNA-seq applications because of their uncertain benefits, concerns about false inferences in downstream applications, and the lack of in-depth benchmark. MethodsHere, we performed two tasks. First, we developed an algorithm using adaptive low-rank full matrix factorization (afMF) based on a previous limited implementation confined to using low rank matrix decomposition (ALRA). Second, to evaluate the impact of various imputation algorithms on downstream analyses, a new benchmark framework incorporating commonly used downstream applications was developed. This benchmark framework put emphasis on real datasets which had ground truth or matched bulk data such that algorithm performance was compared to more convinced data rather than less realistic simulated parameters. ResultsOur results indicated that afMF and ALRA (matrix based) provided good imputation and outperformed raw log-normalization in various downstream applications. afMF outperformed ALRA in several evaluations (cell-level differential expression analysis, GSEA, classification, biomarker prediction, clustering, SC-bulk profiling similarity). Besides, afMF ranked among the top levels in automatic cell type annotation, trajectory inference by DPT, and AUCell & SCENIC. Both showed acceptable scalability, while afMF had longer running time. MAGIC (smoothing based) and AutoClass (deep learning based) also performed well but may produce false positives. In contrast, more complicated methods (other deep learning or model based) were prone to overfitting and data distortion. We also found that certain downstream algorithms are not compatible with imputation, including trajectory inference with Slingshot and cell-cell communication. Prior imputation either showed no improvement or generated false positive findings with these downstream applications. ConclusionsWe hope this in-depth evaluation and the algorithm developed in this study can enhance the selection of appropriate imputation algorithm for specific scRNA-seq downstream analyses. The algorithm and the benchmark framework are available at GitHub: https://github.com/GO3295/SCImputation
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comparison of High-Throughput Single-Cell RNA Sequencing Data Processing Pipelines 96%
- SMNN: Batch Effect Correction for Single-cell RNA-seq data via Supervised Mutual Nearest Neighbor Detection 96%
- SSMD: A semi-supervised approach for a robust cell type identification and deconvolution of mouse transcriptomics data 96%
Similar papers in this journal
Similar papers in this journal
- A systematic evaluation of highly variable gene selection methods for single-cell RNA-sequencing 96%
- Heterogeneous pseudobulk simulation enables realistic benchmarking of cell-type deconvolution methods 96%
- scCDC: a computational method for gene-specific contamination detection and correction in single-cell and single-nucleus RNA-seq data 96%
Similar papers in this journal
- MarcoPolo: a clustering-free approach to the exploration of differentially expressed genes along with group information in single-cell RNA-seq data 96%
- LTMG (Left truncated mixture Gaussian) based modeling of transcriptional regulatory heterogeneities in single cell RNA-seq data - a perspective from the kinetics of mRNA metabolism 96%
- BREM-SC: A Bayesian Random Effects Mixture Model for Joint Clustering Single Cell Multi-omics Data 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.