Back

Mapping of cis-regulatory variants by differential allelic expression analysis identifies candidate risk variants and target genes of 27 breast cancer risk loci

Xavier, J. M.; Magno, R.; Russell, R.; de Almeida, B. P.; Jacinta-Fernandes, A.; Duarte, A.; Dunning, M.; Samarajiwa, S.; O' Reilly, M.; Rocha, C. L.; Rosli, N.; Ponder, B. A. J.; Maia, A.-T.

2022-03-10 genetic and genomic medicine
10.1101/2022.03.08.22271889 medRxiv
Show abstract

Genome-wide association studies (GWAS) have identified hundreds of risk loci for breast cancer, but identifying causal variants and candidate target genes remains challenging. Since most risk loci fall in active gene regulatory regions, we developed a novel approach to identify variants with greater regulatory potential in the diseases tissue of origin. Using genome-wide differential allelic expression (DAE) analysis on microarray data from 64 normal breast tissue samples, we mapped over 54K variants associated with DAE (daeQTLs). We then intersected these with GWAS data to reveal candidate risk regulatory variants and analyzed their cis-acting regulatory potential. We found 122 daeQTLs in 41 loci in active regulatory regions that are in strong linkage disequilibrium with risk-associated variants (risk-daeQTLs). We also identified 65 new candidate target genes in 29 of these loci for which no previous candidates existed. As validation, we identified and functionally characterized five candidate causal variants at the 5q14.1 risk locus targeting the ATG10 and ATP6AP1L genes, likely acting via modulation of alternative transcription and transcription factor binding. Our study demonstrates the power of DAE analysis and daeQTL mapping to understand breast cancer genetic risk, including in complex genetic regulatory landscapes. It additionally provides a genome-wide resource of variants associated with DAE for future functional studies.

Matching journals

The top 12 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.