Back

Data driven refinement of gene expression signatures for enrichment analysis

Wenzel, A. T.; Faraji, F.; Sato, K.; Medetgul-Ernar, K.; Castanza, A.; Sagatelian, R.; Donepudi, G.; Halawa, O.; Wang, J. Y. J.; Gutkind, J. S.; Tamayo, P.; Mesirov, J. P.

2024-11-03 bioinformatics
10.1101/2024.11.03.621768 bioRxiv
Show abstract

Gene set enrichment methods measure biological process or pathway activation in gene expression data by testing coordinate up- or down-regulation of pathway members in a ranked list of genes. These methods rely on curated, annotated gene sets whose members coordinate expression is an indicator of a process or state. We therefore developed the Molecular Signatures Database (MSigDB), a collection of expertly annotated gene sets. While using, enhancing, and expanding MSigDB, we have observed that some gene sets can lack coordinate expression, especially those derived from canonical pathways. To address this challenge, we developed gene set refinement (GSR), a data-driven approach leveraging large-scale multi-omics compendia to extract context-specific sets, deconvolve heterogeneity, and reveal multiple downstream signaling. We applied this method to address cancer biology questions, and demonstrated successful, targeted refinement of existing MSigDB gene sets.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.