WhatIsMyGene: Back to the Basics of Gene Enrichment
Hodge, K. G.; Saethang, T.
Show abstract
WIMG AbstractSince its inception over 25 years ago, gene enrichment has been largely associated with curated gene lists (e.g. GO) that are constructed to represent various biological concepts: the cell cycle, cancer drivers, protein-protein interactions, etc. Researchers expect that a comparison of their own lab-generated lists with curated lists should produce insight. Despite the abundance of such curated lists, we here show that they rarely outperform comparisons against existing individual lab-generated datasets when measured using standard statistical tests of study/study overlap. This demonstration is enabled by the WhatIsMyGene database, which we believe to be the single largest compendium of transcriptomic and micro-RNA perturbation data. The database also houses voluminous proteomic, cell type clustering, lncRNA, and epitranscriptomic data. In the case of enrichment tools that do incorporate specific lab studies in underlying databases, WIMG generally outperforms in the simple task of reflecting back to the user known aspects of the input set (cell type, the type of perturbation, species, etc.), enhancing confidence that unknown aspects of the input may also be revealed in the output. To optimize the ranking of gene enrichment results, WIMG imputes backgrounds to curated gene lists, an approach we investigate with Shannon information analysis. Thus, gene lists that are replete with abundant entities do not inordinately percolate to the highest-ranking positions in output. We delineate a number of other features that should make WIMG indispensable in answering essential questions such as "What processes are embodied in my gene list?" and "What does my gene do?"
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Biology-inspired data-driven quality control for scientific discovery in single-cell transcriptomics 95%
- Robust normalization and transformation techniques for constructing gene coexpression networks from RNA-seq data 95%
- A comparison of marker gene selection methods for single-cell RNA sequencing data 95%
Similar papers in this journal
Similar papers in this journal
- Building, Benchmarking, and Exploring Perturbative Maps of Transcriptional and Morphological Data 96%
- CoVar: A generalizable machine learning approach to identify the coordinated regulators driving variational gene expression 95%
- GeneCOCOA: Detecting context-specific functions of individual genes using co-expression data 94%
Similar papers in this journal
- Multi-view gene panel characterization for spatially resolved omics 95%
- CoRegNet: Unraveling Gene Co-regulation Networks from Public RNA-Seq Repositories Using a Beta-Binomial Statistical Model 94%
- LRcell: detecting the source of differential expression at the sub-cell type level from bulk RNA-seq data 94%
Similar papers in this journal
- Optimal construction of a functional interaction network from pooled library CRISPR fitness screens 94%
- GEOlimma: Differential Expression Analysis and Feature Selection Using Pre-Existing Microarray Data 94%
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.