Deriving Disease Modules from the Compressed Transcriptional Space Embedded in a Deep Auto-encoder
Dwivedi, S. K.; Tjarnberg, A.; Tegner, J.; Gustafsson, M.
Show abstract
Disease modules in molecular interaction maps have been useful for characterizing diseases. Yet biological networks, commonly used to define such modules are incomplete and biased toward some well-studied disease genes. Here we ask whether disease-relevant modules of genes can be discovered without assuming the prior knowledge of a biological network. To this end we train a deep auto-encoder on a large transcriptional data-set. Our hypothesis is that such modules could be discovered in the deep representations within the auto-encoder when trained to capture the variance in the input-output map of the transcriptional profiles. Using a three-layer deep auto-encoder we find a statistically significant enrichment of GWAS relevant genes in the third layer, and to a successively lesser degree in the second and first layers respectively. In contrast, we found an opposite gradient where a modular protein-protein interaction signal was strongest in the first layer but then vanishing smoothly deeper in the network. We conclude that a data-driven discovery approach, without assuming a particular biological network, is sufficient to discover groups of disease-related genes.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Per-sample standardization and asymmetric winsorization lead to accurate clustering of RNA-seq expression profiles 96%
- Non-negative Independent Factor Analysis disentangles discrete and continuous sources of variation in scRNA-seq data 96%
- ECMarker: Interpretable machine learning model identifies gene expression biomarkers predicting clinical outcomes and reveals molecular mechanisms of human disease in early stages 95%
Similar papers in this journal
- scTenifoldNet: a machine learning workflow for constructing and comparing transcriptome-wide gene regulatory networks from single-cell data 95%
- Hierarchical confounder discovery in the experiment-machine learning cycle 94%
- MUSTANG: MUlti-sample Spatial Transcriptomics data ANalysis with cross-sample transcriptional similarity Guidance 94%
Similar papers in this journal
- DOMINO: a novel algorithm for network-based identification of active modules withreduced rate of false calls 95%
- Combinatorial prediction of gene-marker panels from single-cell transcriptomic data. 94%
- KDML: a machine-learning framework for inference of multi-scale gene functions from genetic perturbation screens 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.