Not every gene is special: one simple rule to control the false discovery rate when analysing high-throughput sequencing data
Dos Santos, S. J.; Murariu, A. C.; Silverman, J. D.; Gloor, G. B.
Show abstract
1Differential expression and differential abundance analyses are commonplace in studies employing high-throughput sequencing approaches; however different tools often fail to return comparable results when applied to the same dataset. Most tools employ various normalisations to attempt to correct for technical variation in the count data. Previously, we demonstrated that these normalisations are often inappropriate due to incorrect assumptions regarding the overall scale (i.e. size) of the biological system in question. In this study, we conducted 100 permutation analyses of 10 RNA-seq datasets to show that scale misspecification results in poor control of the false discovery rate by several commonly-used analysis tools. Moreover, we demonstrate that this can be ameliorated by using a scale model in ALDEx2 or ALDEx3 to account for uncertainty around the size of the system scale. We show that there is an inherent trade-off between satisfactory control of false-discovery rates and sensitivity and that no tool offers both. We also quantify how increasing scale uncertainty affects the difference between groups required for a feature to be reported as differentially expressed. Finally, we provide universal guidance on choosing an appropriate amount of scale uncertainty for any type of analysis. Overall, our work highlights the strengths and pitfalls of commonly used tools for differential expression analyses and highlights the choice between sensitivity and false-discovery rate control that all researchers are making when analysing sequencing data.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- miQC: An adaptive probabilistic framework for quality control of single-cell RNA-sequencing data 95%
- XPRESSyourself: Enhancing, Standardizing, and Automating Ribosome Profiling Computational Analyses Yields Improved Insight into Data 94%
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.