Back

Learning the Relationship Between Variants, Metabolic Fluxes and Phenotypes

Kim, D.; Han, S.; Nam, S.-H.; Kim, T. Y.

2024-03-05 bioinformatics
10.1101/2024.03.04.577140 bioRxiv
Show abstract

In the dynamic field of cancer research, the fusion of genetics and machine learning presents a groundbreaking opportunity to understand the intricate links between genomic variations and cancer phenotypes. Despite the wealth of genetic data, translating it into actionable insights remains a challenge, particularly in the context of complex cellular metabolism. VaMP (Variants, Metabolic Fluxes & Phenotypes), our novel approach, addresses this by integrating neural networks and metabolic models, offering a comprehensive framework for deciphering cancer biology. The developed method leverages a novel end-to-end neural network architecture, integrating genome-scale metabolic models (GSMs) to capture the intricate relationship between genetic variations and cellular phenotypes. VaMPs encoder maps genetic variants to metabolic fluxes through a series of carefully designed steps utilizing neural network and GSM, while the decoder predicts phenotype probabilities using the differences between input and reference fluxes. The training data comprise mutation information and phenotypes, eliminating the need for explicit metabolic flux data preparation. Validation experiments on five cancer types demonstrate VaMPs ability to identify significant genes and metabolic signatures. Further utilizing SCREENER and co-occurrence analyses, the assessment reveals VaMPs capacity to anticipate established gene-disease relationships. The identified metabolic signatures are robustly substantiated by diverse literature grounded in experimental studies. Furthermore, an in-depth exploration of the outcomes from five VaMP models involved the correlation analysis for variants and fluxes to establish connections between significant genes and cancer-causing variants. Overall, VaMP can be used as a promising tool for unraveling the complex interplay between genetic alterations and cancer phenotypes, with implications for understanding disease mechanisms and identifying novel therapeutic targets.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.