Machine learning-assisted identification of factors contributing to the technical variability between bulk and single-cell RNA-seq experiments
Lipnitskaya, S.; Shen, Y.; Legewie, S.; Klein, H.; Becker, K.
Show abstract
BackgroundRecent studies in the area of transcriptomics performed on single-cell and population levels reveal noticeable variability in gene expression measurements provided by different RNA sequencing technologies. Due to increased noise and complexity of single-cell RNA-Seq (scRNA-Seq) data over the bulk experiment, there is a substantial number of variably-expressed genes and so-called dropouts, challenging the subsequent computational analysis and potentially leading to false positive discoveries. In order to investigate factors affecting technical variability between RNA sequencing experiments of different technologies, we performed a systematic assessment of single-cell and bulk RNA-Seq data, which have undergone the same pre-processing and sample preparation procedures. ResultsOur analysis indicates that variability between gene expression measurements as well as dropout events are not exclusively caused by biological variability, low expression levels, or random variation. Furthermore, we propose FAVSeq, a machine learning-assisted pipeline for detection of factors contributing to gene expression variability in matched RNA-Seq data provided by two technologies. Based on the analysis of the matched bulk and single-cell dataset, we found the 3-UTR and transcript lengths as the most relevant effectors of the observed variation between RNA-Seq experiments, while the same factors together with cellular compartments were shown to be associated with dropouts. ConclusionsHere, we investigated the sources of variation in RNA-Seq profiles of matched single-cell and bulk experiments. In addition, we proposed the FAVSeq pipeline for analyzing multimodal RNA sequencing data, which allowed to identify factors affecting quantitative difference in gene expression measurements as well as the presence of dropouts. Hereby, the derived knowledge can be employed further in order to improve the interpretation of RNA-Seq data and identify genes that can be affected by assay-based deviations. Source code is available under the MIT license at https://github.com/slipnitskaya/FAVSeq.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Comprehensive benchmark of differential transcript usage analysis for static and dynamic conditions 95%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 95%
- FLYNC: A Machine Learning-Driven Framework for Discovering Long Non-Coding RNAs in Drosophila melanogaster 95%
Similar papers in this journal
- Comparative Analysis of common alignment tools for single cell RNA sequencing 96%
- The case for using Mapped Exonic Non-Duplicate (MEND) read counts in RNA-Seq experiments: examples from pediatric cancer datasets 94%
- Genetic demultiplexing of pooled single-cell RNA-sequencing samples in cancer facilitates effective experimental design 94%
Similar papers in this journal
- On the importance of data transformation for data integration in single-cell RNA sequencing analysis 96%
- Impact of gene annotation choice on the quantification of RNA-seq data 96%
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.