Multi-Ancestry Transcriptome-wide Association Studies Uncover New Insights into Breast Cancer Genetics and Biology
Ping, J.; Jia, G.; Cai, Q.; Guo, X.; Wang, J.; Tao, R.; Li, B.; Bauer, J. A.; Xie, Y.; Ambs, S.; Barnard, M. E.; Chen, Y.; Choi, J.-Y.; Gao, Y.-T.; Garcia-Closas, M.; Gu, J.; Hu, J. J.; Iwasaki, M.; John, E. M.; Kweon, S.-S.; Li, C. I.; Matsuda, K.; Matsuo, K.; Nathanson, K. L.; Nemesure, B.; Olopade, O. I.; Pal, T.; Park, S. K.; Park, B.; Press, M. F.; Sanderson, M.; Sandler, D. P.; Yao, S.; Zheng, Y.; Ahearn, T.; Brewster, A. M.; Falusi, A.; Hennis, A. J.; Ito, H.; Kubo, M.; Lee, E.-S.; Makumbi, T.; Mapoko, B. S.; Noh, D.-Y.; O'Brien, K. M.; Ojengbede, O.; Olshan, A. F.; Park, M.-H.; Reid, S
Show abstract
Genome-wide association studies (GWAS) have identified over 200 genetic risk loci for breast cancer, yet the target genes in these loci remain largely unknown. To address this knowledge gap, we conducted a series of multi-ancestry transcriptome-wide association studies (TWAS) to discover potential breast cancer susceptibility genes. We developed and validated ancestry-specific genetic models to predict levels of gene expression, alternative splicing, and 3 UTR alternative polyadenylation, using genomic and transcriptomic data from normal breast tissue samples of 652 females of African, Asian, or European ancestry. These models were then applied to GWAS data of 178,534 breast cancer cases and 248,300 controls from these ancestry groups for association analyses. We identified 290 genes associated with breast cancer risk, including 103 previously unreported in TWAS and 46 located at least 500Kb away from any previously identified risk variants. Among them, 39 genes exhibited distinct associations with breast cancer risk by estrogen receptor status. The identified genes were enriched in pathways related to homologous recombination, apoptosis, p53, PI3K/AKT/mTOR, estrogen, and IL-2/STAT5 signaling. Single-cell RNA sequencing and in vitro experiment data provided additional functional evidence for 169 genes. Our study uncovered large numbers of candidate breast cancer susceptibility genes and contributed valuable insights into the genetics and biology of this common cancer.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Causal effects of breast cancer risk factors across hormone receptor breast cancer subtypes: A two-sample Mendelian randomization study 95%
- Incorporating alternative Polygenic Risk Scores into the BOADICEA breast cancer risk prediction model 93%
- Cross-cancer genome-wide association study of endometrial cancer and epithelial ovarian cancer identifies genetic risk regions associated with risk of both cancers 92%
Similar papers in this journal
- Pleiotropy-guided transcriptome imputation from normal and tumor tissues identifies new candidate susceptibility genes for breast and ovarian cancer 95%
- Rare coding variants in five DNA damage repair genes associate with timing of natural menopause 93%
- Evaluating Genomic Polygenic Risk Scores for Childhood Acute Lymphoblastic Leukemia in Latinos 92%
Similar papers in this journal
- A joint transcriptome-wide association study across multiple tissues identifies new candidate susceptibility genes for breast cancer 97%
- Segregation analysis of 17,425 population-based breast cancer families: evidence for genetic susceptibility and risk prediction 95%
- A polygenic score-based approach to identify gene-drug interactions stratifying breast cancer risk 93%
Similar papers in this journal
- Allelic expression imbalance of PIK3CA mutations is frequent in breast cancer and prognostically significant 94%
- RNA Sequencing-Based Single Sample Predictors of Molecular Subtype and Risk of Recurrence for Clinical Assessment of Early-Stage Breast Cancer 93%
- Investigating the relationship between breast cancer risk factors and an AI-generated mammographic texture feature in the Nurses' Health Study II 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.