Benchmarking normalisation methods for differential binding analysis in CUT&RUN
Ambani, K.; Ahn, A.; Watt, A.; Blyth, C.; Russ, B.; Taylor, M.; Goel, S.; Oshlack, A.
Show abstract
CUT&RUN (Cleavage Under Targets and Release Using Nuclease) is an increasingly popular method for profiling protein interactions (transcription factors, histone modifications, etc) with DNA across the whole genome. When performing differential binding analysis of CUT&RUN data to identify genomic regions where interaction profiles vary between conditions, data normalisation is essential for accurate biological interpretations. Despite this, there are no clear guidelines on the optimal normalisation method for CUT&RUN datasets. Here, we examine five normalisation approaches (spike-in, library size, background, reads-in-peak and greenlist) and highlight that different methods can result in widely discrepant interpretations of the data. We test these normalisation methods by simulating a variety of plausible differential binding scenarios as well as an in-house generated dataset. We determined that normalisation by either (i) library size or (ii) background to be the most robust. Importantly, we find spike-in normalisation to be the least reliable method. Our findings inform the use of normalisation methods for CUT&RUN data and should thus facilitate reproducible and robust analysis.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Discovery of a non-canonical GRHL1 binding site using deep convolutional and recurrent neural networks 94%
- Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation 93%
- Harnessing changes in open chromatin determined by ATAC-seq to generate insulin-responsive reporter constructs. 93%
Similar papers in this journal
- STARRPeaker: Uniform processing and accurate identification of STARR-seq active regions 95%
- RoAM: computational reconstruction of ancient methylomes and identification of differentially methylated regions 93%
- Positional motif analysis reveals the extent of specificity of protein-RNA interactions observed by CLIP 93%
Similar papers in this journal
- Improved Quality Metrics for Association and Reproducibility in Chromatin Accessibility Data Using Mutual Information 95%
- Comparative analysis of ChIP-exo peak-callers: impact of data quality, read duplication and binding subtypes 95%
- Rescuing Biologically Relevant Consensus Regions Across Replicated Samples 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.