Back

Functional dissection of complex and molecular trait variants at single nucleotide resolution

Siraj, L.; Castro, R. I.; Dewey, H.; Kales, S.; Nguyen, T. T. L.; Kanai, M.; Berenzy, D.; Mouri, K.; Wang, Q.; McCaw, Z. R.; Gosai, S. J.; Aguet, F.; Cui, R.; Vockley, C. M.; Lareau, C. A.; Okada, Y.; Gusev, A.; Jones, T. R.; Lander, E. S.; Sabeti, P. C.; Finucane, H. K.; Reilly, S. K.; Ulirsch, J. C.; Tewhey, R.

2024-05-06 genetics Community evaluation
10.1101/2024.05.05.592437 bioRxiv
Show abstract

Identifying the causal variants and mechanisms that drive complex traits and diseases remains a core problem in human genetics. The majority of these variants have individually weak effects and lie in non-coding gene-regulatory elements where we lack a complete understanding of how single nucleotide alterations modulate transcriptional processes to affect human phenotypes. To address this, we measured the activity of 221,412 trait-associated variants that had been statistically fine-mapped using a Massively Parallel Reporter Assay (MPRA) in 5 diverse cell-types. We show that MPRA is able to discriminate between likely causal variants and controls, identifying 12,025 regulatory variants with high precision. Although the effects of these variants largely agree with orthogonal measures of function, only 69% can plausibly be explained by the disruption of a known transcription factor (TF) binding motif. We dissect the mechanisms of 136 variants using saturation mutagenesis and assign impacted TFs for 91% of variants without a clear canonical mechanism. Finally, we provide evidence that epistasis is prevalent for variants in close proximity and identify multiple functional variants on the same haplotype at a small, but important, subset of trait-associated loci. Overall, our study provides a systematic functional characterization of likely causal common variants underlying complex and molecular human traits, enabling new insights into the regulatory grammar underlying disease risk.

Published in Nature (predicted rank #3) · training set

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.