Back

Whole genome sequencing analysis identifies rare, large-effect non-coding variants and regions associated with circulating protein levels

Hawkes, G.; Chundru, K.; Jackson, L.; Patel, K. A.; Murray, A.; Wood, A. R.; Wright, C. R.; Weedon, M. N.; Frayling, T. M.; Beaumont, R. N.

2023-11-05 genetics
10.1101/2023.11.04.565589 bioRxiv
Show abstract

The role of non-coding rare variation in common phenotypes is largely unknown, due to a lack of whole-genome sequence data, and the difficulty of categorising non-coding variants into biologically meaningful regulatory units. To begin addressing these challenges, we performed a cis association analysis using whole-genome sequence data, consisting of 391 million variants and 1,450 circulating protein levels in [~]20,000 UK Biobank participants. We identified 777 independent rare non-coding single variants associated with circulating protein levels (P<1x10-9), after conditioning on protein-coding and common associated variants. Rare non-coding aggregate testing identified 108 conditionally independent regulatory regions. Unlike protein-coding variation, rare non-coding genetic variation was almost as likely to increase as decrease protein levels. The regions we identified overlapped predicted tissue-specific enhancers more than promoters, suggesting they represent tissue-specific regulatory regions. Our results have important implications for the identification, and role, of rare non-coding variation associated with common human phenotypes.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.