Back

Analyzing and interpreting DNA double-strand break sequencing data

Mitra, A.; Dojer, N.; Fongang, B.; Nde, J.; Zhu, Y.; Rowicka, M.

2020-03-06 bioinformatics
10.1101/2020.03.05.977801 bioRxiv
Show abstract

DNA double-strand breaks (DSBs), are a major threat to genomic stability and may lead to cancer. Several technologies to accurately detect DSBs genome-wide have been developed recently, but still lacking publicly available tools for analysis of the resulting data. Here, we present a step-by-step iSeq package (http://breakome.utmb.edu/software.html), custom designed for analysis and interpretation of DSB-sequencing data. iSeq performs barcode trimming and read counting, and identifies DSB-enriched regions by statistical test and annotate them to the desired genomic features. Applying this package, users can identify and annotate DSB-enriched regions from base pair (eg. Cas9 cleavage sites) up to megabase (eg. DNA replication stress-induced) resolution, and if possible quantify DSB frequencies per cell genome-wide by combining with qDSB-Seq. iSeq can be used for any sequencing-based DSB detection techniques. The analysis for Steps 1-19 can be performed within ~4 hours.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.