Back

DNA-binding factor footprints and enhancer RNAs identify functional non-coding genetic variants

Biddie, S. C.; Hird, E. F.; Weykopf, G.; Friman, E. T.; Bickmore, W. A.

2023-11-20 genomics
10.1101/2023.11.20.567860 bioRxiv
Show abstract

Genome-wide association studies (GWAS) have revealed a multitude of candidate genetic variants affecting the risk of developing complex traits and diseases. However, these highlighted regions are typically in the non-coding genome, and uncovering the functional causative single nucleotide variants (SNVs) is challenging. Prioritisation of variants is commonly based on functional genomic annotation with markers of active regulatory elements, but current approaches still poorly predict functional variants. To address this, we systematically analyse six markers of active regulatory elements for their ability to identify functional variants. We benchmark against molecular quantitative trait loci (molQTL) from assays of regulatory element activity that identify allelic effects on DNA-binding factor occupancy, reporter assay expression, and chromatin accessibility. We identify the combination of DNase footprints and divergent enhancer RNA as markers for functional variants. This signature provides high precision, trading-off low recall, thus substantially reducing candidate variant sets to prioritise variants for functional validation. We present this as a framework called FINDER - Functional SNV IdeNtification using DNase footprints and Enhancer RNA, and demonstrate its utility to prioritise variants using leukocyte count trait and analyse variants in linkage disequilibrium with a lead variant to predict a functional variant in asthma. Our findings have implications for prioritising variants from GWAS, in development of predictive scoring algorithms, and for functionally informed fine mapping approaches.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.