Fast gene set enrichment analysis
Korotkevich, G.; Sukhov, V.; Sergushichev, A.
Show abstract
Gene set enrichment analysis (GSEA) is an ubiquitously used tool for evaluating pathway enrichment in transcriptional data. Typical experimental design consists in comparing two conditions with several replicates using a differential gene expression test followed by preranked GSEA performed against a collection of hundreds and thousands of pathways. However, the reference implementation of this method cannot accurately estimate small P-values, which significantly limits its sensitivity due to multiple hypotheses correction procedure. Here we present FGSEA (Fast Gene Set Enrichment Analysis) method that is able to estimate arbitrarily low GSEA P-values with a high accuracy in a matter of minutes or even seconds. To confirm the accuracy of the method, we also developed an exact algorithm for GSEA P-values calculation for integer gene-level statistics. Using the exact algorithm as a reference we show that FGSEA is able to routinely estimate P-values up to 10-100 with a small and predictable estimation error. We systematically evaluate FGSEA on a collection of 605 datasets and show that FGSEA recovers much more statistically significant pathways compared to other implementations. FGSEA is open source and available as an R package in Bioconductor (http://bioconductor.org/packages/fgsea/) and on GitHub (https://github.com/ctlab/fgsea/).
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Inferring Tumor Progression in Large Datasets 96%
- Learning massive interpretable gene regulatory networks of the human brain by merging Bayesian Networks 96%
- Reconstruction Set Test (RESET): a computationally efficient method for single sample gene set testing based on randomized reduced rank reconstruction error 95%
Similar papers in this journal
- Studying the history of tumor evolution from single-cell sequencing data by exploring the space of binary matrices 95%
- Potpourri: An Epistasis Test Prioritization Algorithm via Diverse SNP Selection 95%
- The statistics of k-mers from a sequence undergoing a simple mutation process without spurious matches 94%
Similar papers in this journal
- Gene prioritization based on random walks with restarts and absorbing states, to define gene sets regulating drug pharmacodynamics from single-cell analyses 95%
- From graph topology to ODE models for gene regulatory networks 94%
- Time Series Experimental Design Under One-Shot Sampling: The Importance of Condition Diversity 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.