Improving the sensitivity of differential-expression analyses for under-powered RNA-seq experiments
Kalinka, A. T.
Show abstract
High-throughput studies, in which thousands of hypothesis tests are conducted simultaneously, can be under-powered when effect sizes are small and there are few replicates. Here, I describe an approach to estimate the FDR for a given experiment such that the ground truth is known. A decision boundary between true and false positive calls can then be learned from the data itself along the axes of fold change and expression level. By excluding hits that fall into the false positive space, the FDR of any given method can be controlled providing a means to employ less conservative methods for detecting differential expression without incurring the usual loss of precision. I show that coupling this approach with a feature-selection method - an elastic-net logistic regression - can increase sensitivity 10-fold above what is achievable with the prevailing methods of the day. An R package implementing these methods is available at https://github.com/alextkalinka/delboy.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Dividing out quantification uncertainty enables assessment of differential transcript usage with limma and edgeR 94%
- Matrix factorization and transfer learning uncover regulatory biology across multiple single-cell ATAC-seq data sets 94%
- Dividing out quantification uncertainty allows efficient assessment of differential transcript expression with edgeR 94%
Similar papers in this journal
- SPECK: An Unsupervised Learning Approach for Cell Surface Receptor Abundance Estimation for Single Cell RNA-Sequencing Data 93%
- WAS IT A MATch I SAW? Approximate palindromes lead to overstated false match rates in benchmarks using reversed sequences 92%
- Single cell gene set scoring with nearest neighborgraph smoothed data (gssnng). 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.