Greenscreen decreases Type I Errors and increases true peak detection in genomic datasets including ChIP-seq
Klasfeld, S.; Wagner, D.
Show abstract
Chromatin immunoprecipitation followed by sequencing (ChIP-seq) is used widely to identify both factor binding to genomic DNA and chromatin modifications. Analysis of ChIP-seq data is impacted by regions of the genome which generate ultra-high artifactual signals. To remove these signals from ChIP-seq data, ENCODE developed blacklists, comprehensive sets of regions defined by low mappability and ultra-high signals for human, mouse, worm, and flies. Currently, blacklists are not available for many model and non-model species. Here we describe an alternative approach for removing false-positive peaks we called "greenscreen". Greenscreen is facile to implement, requires few input samples, and uses analysis tools frequently employed for ChIP-seq. We show that greenscreen removes artifact signal as effectively as blacklists in Arabidopsis and human ChIP-seq datasets while covering less of the genome, dramatically improving ChIP-seq data quality. Greenscreen filtering reveals true factor binding overlap and of occupancy changes in different genetic backgrounds or tissues. Because it is effective with as few as three inputs, greenscreen is readily adaptable for use in any species or genome build. Although developed for ChIP-seq, greenscreen also identifies artifact signals from other genomic datasets including CUT&RUN. Finally, we present an improved ChIP-seq pipeline which incorporates greenscreen, that detects more true peaks than published methods. One Sentence SummaryA facile method for removing artifact signal from ChIP-seq that improves downstream analyses
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Major cell-types in multiomic single-nucleus datasets impact statistical modeling of links between regulatory sequences and target genes 94%
- Uncovering uncharacterized binding of transcription factors from ATAC-seq footprinting data 94%
- Fast Fourier Transform is a training-free, ultrafast, highly efficient, and fully interpretable approach for epigenomic data compression 94%
Similar papers in this journal
- Epigenome and interactome profiling uncovers principles of distal regulation in the barley genome 95%
- Whole genome functional characterization of RE1 silencers using a modified massively parallel reporter assay 94%
- Non-additive genetic components contribute significantly to population-wide gene expression variation 94%
Similar papers in this journal
- SCENIC+: single-cell multiomic inference of enhancers and gene regulatory networks 95%
- DeepC: Predicting chromatin interactions using megabase scaled deep neural networks and transfer learning. 95%
- Simultaneous profiling of chromatin accessibility and methylation on human cell lines with nanopore sequencing 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.