Back

A-LAVA: Detecting impact of germline variants on metabolic pathways in cancer genomes

Jalilkhany, M.; Gallagher, K.; Wehner, I.; Hamilton, P.; MacPherson, S.; McPhedran, S.; Nathoo, F.; Lum, J.; Numanagic, I.

2025-09-04 bioinformatics
10.1101/2025.08.30.673276 bioRxiv
Show abstract

The metabolic landscape of cancer has been widely studied, especially in the context of somatic mutations. However, the impact of inherited germline variants upon the metabolic genes interaction still remains unexplored. In this work, we present a computational pipeline named A-LAVA for the detection and analysis of germline variants that affect metabolic pathways in cancer. Our pipeline enables analysis at three different levels: SNP, gene, and pathway-based analysis. The first steps consist of detecting statistically significant SNPs through standardized GWAS pipelines and, subsequently, genes associated with metabolic traits through gene-level analysis. Then, A-LAVA performs gene set analysis (GSA) to further explore the effect of detected associations on metabolic pathways. This analysis is done through a statistical model that newly corrects for the confounding effects arising from overlapping gene sets, in addition to other corrections performed by the current best practices. Our analysis conducted on TCGA data shows that SNP and gene-level results identified key associations and that A-LAVAs GSA approach improved the overall accuracy both on synthetic and real data by correctly correcting for overlapping genes, refining significance thresholds, and reducing false positives, thus leading to more reliable metabolic pathway rankings and a more robust framework for gene set analysis. CCS CONCEPTSO_LIApplied computing [->]Bioinformatics; Metabolomics / metabonomics. C_LI

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.