Back

An Expanded Registry of Candidate cis-Regulatory Elements for Studying Transcriptional Regulation

Moore, J. E.; Pratt, H. E.; Fan, K.; Phalke, N.; Fisher, J.; Elhajjajy, S. I.; Andrews, G.; Gao, M.; Shedd, N.; Fu, Y.; Lacadie, M. C.; Meza, J.; Ganna, M.; Choudhury, E.; Swofford, R.; Farrell, N. P.; Pampari, A.; Ramalingam, V.; Reese, F.; Borsari, B.; Yu, X.; Wattenberg, E. S.; Ruiz-Romero, M.; Razavi-Mohseni, M.; Xu, J.; Galeev, T.; Beer, M. A.; Guigo, R.; Gerstein, M.; Engreitz, J. M.; Ljungman, M.; Reddy, T. E.; Snyder, M.; Epstein, C. B.; Gaskell, E.; Bernstein, B. E.; Dickel, D. E.; Visel, A.; Pennacchio, L. A.; Mortazavi, A.; Kundaje, A.; Weng, Z.

2024-12-26 genomics
10.1101/2024.12.26.629296 bioRxiv
Show abstract

Mammalian genomes contain millions of regulatory elements that control the complex patterns of gene expression. Previously, The ENCODE consortium mapped biochemical signals across many cell types and tissues and integrated these data to develop a Registry of 0.9 million human and 300 thousand mouse candidate cis-Regulatory Elements (cCREs) annotated with potential functions1. We have expanded the Registry to include 2.35 million human and 927 thousand mouse cCREs, leveraging new ENCODE datasets and enhanced computational methods. This expanded Registry covers hundreds of unique cell and tissue types, providing a comprehensive understanding of gene regulation. Functional characterization data from assays like STARR-seq, MPRA, CRISPR perturbation, and transgenic mouse assays now cover over 90% of human cCREs, revealing complex regulatory functions. We identified thousands of novel silencer cCREs and demonstrated their dual enhancer/silencer roles in different cellular contexts. Integrating the Registry with other ENCODE annotations facilitates genetic variation interpretation and trait-associated gene identification, exemplified by discovering KLF1 as a novel causal gene for red blood cell traits. This expanded Registry is a valuable resource for studying the regulatory genome and its impact on health and disease.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.