Built on sand: the shaky foundations of simulating single-cell RNA sequencing data
Crowell, H. L.; Leonardo, S. X. M.; Soneson, C.; Robinson, M. D.
Show abstract
With the emergence of hundreds of single-cell RNA-sequencing (scRNA-seq) datasets, the number of computational tools to analyse aspects of the generated data has grown rapidly. As a result, there is a recurring need to demonstrate whether newly developed methods are truly performant - on their own as well as in comparison to existing tools. Benchmark studies aim to consolidate the space of available methods for a given task, and often use simulated data that provide a ground truth for evaluations. Thus, demanding a high quality standard for synthetically generated data is critical to make simulation study results credible and transferable to real data. Here, we evaluated methods for synthetic scRNA-seq data generation in their ability to mimic experimental data. Besides comparing gene- and cell-level quality control summaries in both one- and two-dimensional settings, we further quantified these at the batch- and cluster-level. Secondly, we investigate the effect of simulators on clustering and batch correction method comparisons, and, thirdly, which and to what extent quality control summaries can capture reference-simulation similarity. Our results suggest that most simulators are unable to accommodate complex designs without introducing artificial effects; they yield over-optimistic performance of integration, and potentially unreliable ranking of clustering methods; and, it is generally unknown which summaries are important to ensure effective simulation-based method comparisons.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 97%
- A comparison of marker gene selection methods for single-cell RNA sequencing data 97%
- Robustness and applicability of functional genomics tools on scRNA-seq data 97%
Similar papers in this journal
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 95%
- Optimal tuning of weighted kNN- and diffusion-based methods for denoising single cell genomics data 95%
- HiCImpute: A Bayesian Hierarchical Model for Identifying Structural Zeros and Enhancing Single Cell Hi-C Data. 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.