The CUT&RUN Blacklist of Problematic Regions of the Genome
Nordin, A.; Zambanini, G.; Pagella, P.; Cantu, C.
Show abstract
Cleavage Under Targets and Release Using Nuclease (CUT&RUN) is an increasingly popular technique to map genome-wide binding profiles of histone modifications, transcription factors and co-factors. The ENCODE project and others have compiled blacklists for ChIP-seq which have been widely adopted: these lists contain regions of high and unstructured signal, regardless of cell type or protein target. While CUT&RUN obtains similar results to ChIP-seq, its biochemistry and subsequent data analyses are different. We found that this results in a CUT&RUN-specific set of undesired high-signal regions. For this reason, we have compiled blacklists based on CUT&RUN data for the human and mouse genomes, identifying regions consistently called as peaks in negative controls by the CUT&RUN peak caller SEACR. Using published CUT&RUN data from our and other labs, we show that the CUT&RUN blacklist regions can persist even when peak calling is performed with SEACR against a negative control, and after ENCODE blacklist removal. Moreover, we experimentally validated the CUT&RUN Blacklists by performing reiterative negative control experiments in which no specific protein is targeted, showing that they capture >80% of the peaks identified. We propose that removing these problematic regions prior to peak calling can substantially improve the performance of SEACR-based peak calling in CUT&RUN experiments, resulting in more reliable peak datasets.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RADAR: Differential analysis of MeRIP-seq data with a random effect model 95%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 94%
- JAFFAL: Detecting fusion genes with long read transcriptome sequencing 94%
Similar papers in this journal
- MUFFIN : A suite of tools for the analysis of functional sequencing data 96%
- Kmerator Suite: design of specific k-mer signatures andautomatic metadata discovery in large RNA-Seq datasets. 94%
- Towards Personalized Epigenomics: Learning Shared Chromatin Landscapes and Joint De-Noising of Histone Modification Assays 93%
Similar papers in this journal
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 96%
- Rescuing Biologically Relevant Consensus Regions Across Replicated Samples 96%
- ChIPbinner: An R package for analyzing broad histone marks binned in uniform windows from ChIP-Seq or CUT&RUN/TAG data 95%
Similar papers in this journal
- Copy number normalization distinguishes differential signals driven by copy number differences in ATAC-seq and ChIP-seq 96%
- Characterization of a strain-specific CD-1 reference genome reveals potential inter- and intra-strain functional variability 94%
- Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.