Back

Generative neural networks separate common and specific transcriptional responses

Lee, A. J.; Mould, D. L.; Crawford, J.; Hu, D.; Powers, R. K.; Doing, G.; Costello, J. C.; Hogan, D. A.; Greene, C. S.

2021-05-24 bioinformatics
10.1101/2021.05.24.445440 bioRxiv
Show abstract

Genome-wide transcriptome profiling identifies genes that are prone to differential expression across contexts ("common DEGs"), as well as genes with changes specific to the experimental manipulation. Distinguishing common DEGs from those that are specifically changed in a context of interest allows more efficient prediction of which genes are specific to a given biological process under scrutiny. Currently, commonly differentially expressed genes or pathways can only be identified through the laborious manual curation of highly controlled experiments, an inordinately time-consuming and impractical endeavor. Here we pioneer an approach for identifying common patterns using generative neural networks. This approach produces a background set of transcriptomic experiments from which a null distribution of gene and pathway changes can be generated. By comparing the set of differentially expressed genes found in a target experiment against the generated background set, common results can be easily separated from specific ones. This "Specific cOntext Pattern Highlighting In Expression data" (SOPHIE) approach is broadly applicable to new platforms or any species with a large collection of gene expression data. We apply SOPHIE to diverse datasets including those from human, human cancer, and the bacterial pathogen Pseudomonas aeruginosa. SOPHIE identifies common DEGs in concordance with previously described, manually and systematically determined common DEGs. Further, molecular validation indicates that SOPHIE detects highly specific, but low magnitude, biologically relevant, transcriptional changes. SOPHIEs measure of specificity can complement log fold change values generated from traditional differential expression analyses. For example, by filtering the set of differentially expressed genes, one can identify those genes that are specifically relevant to the experimental condition of interest. Consequently, these results can inform future research directions.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.