Benchmarking scRNA-seq imputation tools with respect to network inference highlights deficits in performance at high levels of sparsity
Steinheuer, L. M.; Canzler, S.; Hackermüller, J.
Show abstract
Gene correlation network inference from single-cell transcriptomics data potentially allows to gain unprecendented insights into cell type-specific regulatory programs. ScRNA-seq data is severely affected by dropout, which significantly hampers and restrains current downstream analysis. Although newly developed tools are capable to deal with sparse data, no appropriate single-cell network inference workflow has been established. A potential way to end this deadlock is the application of data imputation methods, which already proofed to be useful in specific contexts of single-cell data analysis, e.g., recovering cell clusters. In order to infer cell-type specific networks, two prerequisites must be met: the identification of cluster-specific cell-types and the network inference itself. Here, we propose a benchmarking framework to investigate both objections. By using suitable reference data with inherent correlation structure, six representative imputation tools and appropriate evaluation measures, we were able to systematically infer the impact of data imputation on network inference. Major network structures were found to be preserved in low dropout data sets. For moderately sparse data sets, DCA was able to recover gene correlation structures, although systematically introducing higher correlation values. No imputation tool was able to recover true signals from high dropout data. However, by using an additional biological data set we could show that cell-cell correlation by means of specific marker gene expression was not compromised through data imputation. Our analysis showed that network inference is feasible for low and moderately sparse data sets by using the unimputed and DCA-prepared data, respectively. High sparsity data, on the other side, still pose a major problem since current imputation techniques are not able to facilitate network inference. The annotation of cluster-specific cell-types as a prerequisite is not hampered by data imputation but their power to restore the deeply hidden correlation structures is still not sufficient enough.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Gene regulation network inference using k-nearest neighbor-based mutual information estimation- Revisiting an old DREAM 96%
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 96%
- SUBATOMIC: a SUbgraph BAsed mulTi-OMIcs Clustering framework to analyze integrated multi-edge networks 96%
Similar papers in this journal
- G2S3: a gene graph-based imputation method for single-cell RNA sequencing data 96%
- HELP: A computational framework for labelling and predicting human common and context-specific essential genes 96%
- Mcadet: a feature selection method for fine-resolution single-cell RNA-seq data based on multiple correspondence analysis and community detection 96%
Similar papers in this journal
- Comparative Analysis of common alignment tools for single cell RNA sequencing 96%
- Stardust: improving spatial transcriptomics data analysis through space aware modularity optimization based clustering. 95%
- scShapes: A statistical framework for identifying distribution shapes in single-cell RNA-sequencing data 94%
Similar papers in this journal
- Perturbation-based gene regulatory network inference to unravel oncogenic mechanisms 96%
- Naturally occurring combinations of receptors from single cell transcriptomics in endothelial cells 95%
- Finding disease modules for cancer and COVID-19 in gene co-expression networks with the Core&Peel method 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.