Organ-specific prioritization and annotation of non-coding regulatory variants in the human genome
Zhao, N.; Dong, S.; Boyle, A. P.
Show abstract
Identifying non-coding regulatory variants in the human genome remains a challenging task in genomics. Recently, we released the second version of our leading regulatory variant database, RegulomeDB. Building upon this comprehensive database, we developed a novel machine-learning architecture, TLand, which utilizes RegulomeDB-derived features to predict regulatory variants at the cell- or organ-specific level. In our holdout benchmarking, TLand consistently outperformed state-of-the-art models, demonstrating its ability to generalize to new cell lines or organs. We trained three types of organ-specific TLand models to overcome the common model bias toward high data availability cell lines or organs. These models accurately prioritize relevant organs for 2 million GWAS SNPs associated with GWAS traits. Moreover, our analysis of top-scoring variants in specific organ models showed a high enrichment of relevant GWAS traits. We expect that TLand and RegulomeDB will further advance our ability to understand human regulatory variants genome-wide.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromatin information content landscapes inform transcription factor and DNA interactions 95%
- The molecular basis, genetic control and pleiotropic effects of local gene co-expression 94%
- Leveraging supervised learning for functionally-informed fine-mapping of cis-eQTLs identifies an additional 20,913 putative causal eQTLs 94%
Similar papers in this journal
Similar papers in this journal
- Specifying cellular context of transcription factor regulons for exploring context-specific gene regulation programs 96%
- Single-cell reference mapping to construct and extend cell type hierarchies 94%
- Comparative single-cell transcriptomic analysis reveals putative differentiation drivers and potential origin of vertebrate retina 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.