Annotating and prioritizing human non-coding variants with RegulomeDB
Dong, S.; Zhao, N.; Spragins, E.; Kagda, M. S.; Li, M.; Assis, P. R.; Jolanki, O.; Luo, Y.; Cherry, J. M.; Boyle, A. P.; Hitz, B. C.
Show abstract
Nearly 90% of the disease risk-associated variants identified from genome-wide association studies (GWAS) are in non-coding regions of the genome. The annotations obtained from analyzing functional genomics assays can provide additional information to pinpoint causal variants, which are often not the lead variants identified from association studies. However, the lack of available annotation tools limits the use of such data. To address the challenge, we have previously built the RegulomeDB database for prioritizing and annotating variants in non-coding regions1, which has been a highly utilized resource for the research community (Supplementary Fig. 1). RegulomeDB annotates a variant by intersecting its position with genomic intervals identified from functional genomic assays and computational approaches. It also incorporates those hits of a variant into a heuristic ranking score, representing its potential to be functional in regulatory elements. Here we present a newer version of the RegulomeDB web server, RegulomeDB v2.1 (http://regulomedb.org). We improve and boost annotation power by incorporating thousands of newly processed data from functional genomic assays in GRCh38 assembly, and now include probabilistic scores from the SURF algorithm that was the top performing non-coding variant predictor in CAGI 52. We also provide interactive charts and genome browser views to allow users an easy way to perform exploratory analyses in different tissue contexts.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Central role of glycosylation processes in human genetic susceptibility to SARS-CoV-2 infections with Omicron variants 94%
- Large scale genome-wide association study in a Japanese population identified 45 novel susceptibility loci for 22 diseases 94%
- Scalable generalized linear mixed model for region-based association tests in large biobanks and cohorts 94%
Similar papers in this journal
Similar papers in this journal
- Bayesian model comparison for rare variant association studies 94%
- Integration of genetic fine-mapping and multi-omics data reveals candidate effector genes for hypertension 94%
- Widespread recessive effects on common diseases in a cohort of 44,000 British Pakistanis and Bangladeshis with high autozygosity 94%
Similar papers in this journal
- Genome-wide prediction of pathogenic gain- and loss-of-function variants from ensemble learning of diverse feature set 94%
- A systematic analysis of splicing variants identifies new diagnoses in the 100,000 Genomes Project. 94%
- MetaRNN: Differentiating Rare Pathogenic and Rare Benign Missense SNVs and InDels Using Deep Learning 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.