Back

Isoform-level analyses of 6 cancers uncover extensive genetic risk mechanisms undetected at the gene-level

Chang, Y.-H.; Head, S. T.; Harrison, T.; Yu, Y.; Huff, C. D.; Pasaniuc, B.; Lindstroem, S.; Bhattacharya, A.

2024-10-30 genetic and genomic medicine
10.1101/2024.10.29.24316388 medRxiv
Show abstract

Integrating genome-wide association study (GWAS) and transcriptomic datasets can help identify mediators for genetic risk of cancer. Traditional methods often are insufficient as they rely on total gene expression measures and overlook alternative splicing, which generates different transcript-isoforms with potentially distinct effects. Here, we integrate multi-tissue isoform expression data from the Genotype Tissue-Expression Project with GWAS summary statistics (all N > 20,000 cases) to identify isoform- and gene-level associations with six cancers (breast, endometrial, colorectal, lung, ovarian, prostate) and six related cancer subtype classifications (N = 12 total). Directly modeling isoforms using transcriptome-wide association studies (isoTWAS) significantly improves discovery of genetic associations compared to gene-level approaches, identifying 164% more significant associations (6,163 vs. 2,336) with isoTWAS-prioritized genes enriched 4-fold for evolutionarily-constrained genes. isoTWAS tags transcriptomic associations at 52% more independent GWAS loci across the six cancers. Isoform expression mediates an estimated 63% greater proportion of cancer risk SNP heritability compared to gene expression. We highlight several isoTWAS associations that demonstrate GWAS colocalization at the isoform level but not at the gene level, including CLPTM1L (lung cancer), LAMC1 (colorectal), and BABAM1 (breast). These results underscore the importance of modeling isoforms to maximize discovery of genetic risk mechanisms for cancers.

Published in British Journal of Cancer · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.