HiCUP-Plus: a fast open-source pipeline for accurately processing large scale Hi-C sequence data
Kelly, S. T.; Yuhara, S.
Show abstract
Hi-C is an unbiased genome-wide assay to study 3D chromosome conformation and gene-regulation. The HiCUP pipeline is an open-source tool to process Hi-C from massively parallel sequencing while accounting for biases specific to the restriction enzyme digests used. It is an excellent solution tailored to analyse this technique, however the latest aligner supported by the current release is Bowtie2. To improve the computational performance and mapping accuracy when using the HiCUP pipeline, we have modified it to optionally call the HiSAT2 and Dragen aligners. This allows using the HiCUP pipeline with 3rd party aligners, including the commercially-licensed high performance Dragen aligner. The HiCUP+ pipeline is modified extensively to be compatible with Dragen outputs while ensuring that the same results as the original pipeline can be reproduced with the Bowtie or Bowtie2 aligners. Using the highly accurate HiSAT2 or Dragen aligners produces larger outputs with a higher proportion of uniquely mapped read pairs. It is therefore feasible to leverage the reduced compute-time of Dragen to reduce compute costs and turnaround-time without compromising quality of results. The HiCUP pipeline and Dragen both compute rich summary information.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- ILIAD: A suite of automated Snakemake workflows for processing genomic data for downstream applications 96%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 96%
- SequelTools: A Suite of Tools for Working with PacBio Sequel Raw Sequence Data 96%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.