Low-dimensional representations of genome-scale metabolism
Cain, S.; Merzbacher, C.; Oyarzun, D. A.
Show abstract
Cellular metabolism is a highly interconnected network with thousands of reactions that convert nutrients into the molecular building blocks of life. Metabolic connectivity varies greatly with cellular context and environmental conditions, and it remains a challenge to compare genome-scale metabolism across cell types because of the high dimensionality of the reaction flux space. Here, we employ self-supervised learning and genome-scale metabolic models to compress the flux space into low-dimensional representations that preserve structure across cell types. We trained variational autoencoders (VAEs) on large fluxomic data (N = 800, 000) sampled from patient-derived models for various cancer cell types. The VAE embeddings have an improved ability to distinguish cell types than the uncompressed fluxomic data, and sufficient predictive power to classify cell types with high accuracy. We tested the ability of these classifiers to assign cell type identities to unlabelled patient-derived metabolic models not employed during VAE training. We further employed the pre-trained VAE to embed another 38 cell types and trained multilabel classifiers that display promising generalization performance. Our approach distils the metabolic space into a semantically rich vector that can be used as a foundation for predictive modelling, clustering or comparing metabolic capabilities across organisms.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- scConfluence : single-cell diagonal integration with regularized Inverse Optimal Transport on weakly connected features 96%
- OmicVerse: A single pipeline for exploring the entire transcriptome universe 96%
- Triple-effect correction for Cell Painting data with contrastive and domain-adversarial learning 96%
Similar papers in this journal
- Reconstructing Kinetic Models for Dynamical Studies of Metabolism using Generative Adversarial Networks 96%
- Delineating the Effective Use of Self-Supervised Learning in Single-Cell Genomics 96%
- Gene set inference from single-cell sequencing data using a hybrid of matrix factorization and variational autoencoders 96%
Similar papers in this journal
- Clustering-independent estimation of cell abundances in bulk tissues using single-cell RNA-seq data 95%
- Quantifying tumor specificity using Bayesian probabilistic modeling for drug target discovery and prioritization 94%
- Single-cell multi-omic topic embedding reveals cell-type-specific and COVID-19 severity-related immune signatures 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.