Back

Multi-center integrated analysis of non-coding CRISPR screens

Yao, D.; Tycko, J.; Oh, W.; Bounds, L. R.; Gosai, S. J.; Lataniotis, L.; Mackay-Smith, A.; Doughty, B. R.; Gabdank, I.; Schmidt, H.; Youngworth, I.; Andreeva, K.; Ren, X.; Barrera, A.; Luo, Y.; Siklenka, K.; Yardimci, G. G.; The ENCODE4 Consortium, ; Tewhey, R.; Kundaje, A.; Greenleaf, W. J.; Sabeti, P. C.; Leslie, C.; Pritykin, Y.; Moore, J. E.; Beer, M. A.; Gersbach, C.; Reddy, T. E.; Shen, Y.; Engreitz, J. M.; Bassik, M. C.; Reilly, S. K.

2022-12-22 genomics
10.1101/2022.12.21.520137 bioRxiv
Show abstract

The ENCODE Consortiums efforts to annotate non-coding, cis-regulatory elements (CREs) have advanced our understanding of gene regulatory landscapes which play a major role in health and disease. Pooled, non-coding CRISPR screens are a promising approach for systematically investigating gene regulatory mechanisms. Here, the ENCODE Functional Characterization Centers report 109 screens comprising 346,970 individual perturbations across 13.3Mb of the genome, using a variety of methods, readouts, and statistical analyses. Across 332 functionally confirmed CRE-gene links, we identify principles for screening endogenous, non-coding elements for causal regulatory mechanisms. Nearly all CREs show strong evidence of open chromatin, and targeting accessibility peak summits is a critical component of our proposed sgRNA design rules. We provide experimental guidelines to accurately detect CREs with variable, often low, transcriptional effects. We discover a previously undescribed DNA strand-bias for CRISPRi in transcribed regions with implications for screen design and analysis. Benchmarking five screen analysis tools, we find CASA produces the most conservative CRE calls and is robust to artifacts of low-specificity sgRNAs. Together, we provide an accessible data resource, predesigned sgRNAs targeting 3,275,697 ENCODE SCREEN candidate CREs, and screening guidelines to accelerate functional characterization of the non-coding genome.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.