Representational Learning from Healthy Multi-Tissue Human RNA-seq Data such that Latent Space Arithmetics Extracts Disease Modules
de Weerd, H. A.; Guala, D.; Gustafsson, M.; Synnergren, J.; Tegner, J.; Lubovac-Pilav, Z.; Magnusson, R.
Show abstract
1Developing computational analyses of transcriptomic data has dramatically improved our understanding of complex multifactorial diseases. However, such approaches are limited to small sample sets of disease-affected material, thus being sensitive to statistical biases and noise. Here, we ask if a variational autoencoder (VAE) trained on large groups of healthy, human RNA-seq data of multiple tissues can capture the fundamental healthy gene regulation system such that the learned representation generalizes to account for unseen disease changes. To this end, we trained a multi-scale representation to encode cellular processes ranging from cell types to genegene interactions. Importantly, we found that the learned healthy representations could predict unseen gene expression changes from 25 independent disease datasets. We extracted and decoded disease-specific signals from the VAE latent space to dissect this finding. Interestingly, the gene modules corresponding to this signal contained more disease-specific genes than the respective differential expression analysis in 20 of 25 cases. Finally, we matched genes related to the disease signals to known drug targets. We could extract sets of known and potential pharmaceutical candidates from this analysis and demonstrate the utility in three use cases. In summary, our study showcases how data-driven representation learning using a VAE as a foundational model allows an arithmetic deconstruction of the latent space such that biological insights enable the dissection of disease mechanisms and drug targets. Our model is available at https://github.com/ddeweerd/VAE_Transcriptomics/.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scTenifoldNet: a machine learning workflow for constructing and comparing transcriptome-wide gene regulatory networks from single-cell data 96%
- scELMo: Embeddings from Language Models are Good Learners for Single-cell Data Analysis 95%
- Bi-level Graph Learning Unveils Prognosis-Relevant Tumor Microenvironment Patterns in Breast Multiplexed Digital Pathology 95%
Similar papers in this journal
Similar papers in this journal
- FastCCC: A permutation-free framework for scalable, robust, and reference-based cell-cell communication analysis in single cell transcriptomics studies 96%
- Learning interpretable cellular and gene signature embeddings from single-cell transcriptomic data 96%
- Projecting genetic associations through gene expression patterns highlights disease etiology and drug mechanisms 96%
Similar papers in this journal
- scCross: A Deep Generative Model for Unifying Single-cell Multi-omics with Seamless Integration, Cross-modal Generation, and In-silico Exploration 96%
- CMOT: Cross Modality Optimal Transport for multimodal inference 96%
- Integrating temporal single-cell gene expression modalities for trajectory inference and disease prediction 96%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.