Prioritizing context-specific genetic risk mechanisms in 11 solid cancers
Wu, X.; Kim, A.; Breeze, C. E.; O'Mara, T. A.; Ramachandran, D.; Dork, T.; Koutros, S.; Rothman, N.; Prokunina-Olsson, L.; Mancuso, N.; Lindstroem, S.; Kraft, P.
Show abstract
BackgroundWhile genome-wide association studies (GWAS) have identified hundreds of cancer-associated genetic variants, the specific biological contexts where these variants exert their effects remain largely unknown. We aimed to prioritize context-specific genetic risk mechanisms for 11 solid cancers at both genome-wide and single-variant resolutions. MethodsWe integrated cancer GWAS summary statistics from European ancestry samples (avg. n cases=47,856) with [~]1,500 context-specific annotations representing candidate cis-regulatory elements. For genome-wide analysis, we applied CT-FM, a method that leverages heritability enrichment estimates and an annotation correlation matrix to select likely disease-relevant biological contexts. After identifying putative causal SNPs (PIP[≥]0.5) via functionally informed fine-mapping, we used CT-FM-SNP to identify relevant contexts for individual variants. A combined SNP-to-gene framework was applied to construct putative {regulatory SNP-context-gene-cancer} quadruplets. ResultsStratified LD score regression analysis identified 52 annotations with significant heritability enrichment (Bonferroni-corrected P[≤]0.05). CT-FM prioritized four high-confidence (PIP[≥]0.5) biological contexts: mammary luminal epithelial cells for breast cancer, a prostate cancer epithelial cell line (VCaP) for prostate cancer, and bulk tumor tissue contexts for colorectal and renal cancers. Variant-level analysis of hundreds of putatively causal SNPs corroborated these findings and identified additional high-confidence contexts for other malignancies, including estrogen receptor-negative breast cancer and bladder cancer. A total of 489 putative regulatory quadruplets were constructed, proposing specific molecular mechanisms underlying the observed GWAS signals. ConclusionThese findings advance our understanding of genetic susceptibility to different cancers. Future work in larger, more diverse GWAS, coupled with more comprehensive annotation atlases, is essential to expand upon and validate our results.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A joint transcriptome-wide association study across multiple tissues identifies new candidate susceptibility genes for breast cancer 96%
- A polygenic score-based approach to identify gene-drug interactions stratifying breast cancer risk 95%
- The contribution of coding variants to the heritability of multiple cancer types using UK Biobank whole-exome sequencing data 95%
Similar papers in this journal
- Genome-wide association study identifies 32 novel breast cancer susceptibility loci from overall and subtype-specific analyses 96%
- Large scale genome-wide association study in a Japanese population identified 45 novel susceptibility loci for 22 diseases 95%
- Genomic evolution of pancreatic cancer at single-cell resolution 94%
Similar papers in this journal
- Pan-cancer analysis demonstrates that integrating polygenic risk scores with modifiable risk factors improves risk prediction 96%
- Cross-dataset pan-cancer detection: Correlating cell-free DNA fragment coverage with open chromatin sites across cell types 96%
- Risk factors for eight common cancers revealed from a phenome-wide Mendelian randomisation analysis of 378,142 cases and 485,715 controls 96%
Similar papers in this journal
Similar papers in this journal
- Causal effects of breast cancer risk factors across hormone receptor breast cancer subtypes: A two-sample Mendelian randomization study 92%
- A new colorectal cancer risk prediction model incorporating family history, personal and environmental factors 92%
- Genetic analysis of functional rare germline variants across 9 cancer types from the DiscovEHR study 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.