Probabilistic method corrects previously uncharacterized Hi-C artifact
Shen, Y.; Kingsford, C.
Show abstract
Three-dimensional chromosomal structure plays an important role in gene regulation. Chromosome conformation capture techniques, especially the high-throughput, sequencing-based technique Hi-C, provide new insights on spatial architectures of chromosomes. However, Hi-C data contains artifacts and systemic biases that substantially influence subsequent analysis. Computational models have been developed to address these biases explicitly, however, it is difficult to enumerate and eliminate all the biases in models. Other models are designed to correct biases implicitly, but they will also be invalid in some situations such as copy number variations. We characterize a new kind of artifact in Hi-C data. We find that this artifact is caused by incorrect alignment of Hi-C reads against approximate repeat regions and can lead to erroneous chromatin contact signals. The artifact cannot be corrected by current Hi-C correction methods. We design a probabilistic method and develop a new Hi-C processing pipeline by integrating our probabilistic method with the HiC-Pro pipeline. We find that the new pipeline can remove this new artifact effectively, while preserving important features of the original Hi-C matrices.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 96%
- Determination of complete chromosomal haplotypes by bulk DNA sequencing 95%
- A unified encyclopedia of human functional DNA elements through fully automated annotation of 164 human cell types 95%
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- A map of cis-regulatory modules and constituent transcription factor binding sites in 80% of the mouse genome 94%
- XCVATR: Detection and Characterization of Variant Impact on the Embeddings of Single -Cell and Bulk RNA-Sequencing Samples 94%
- SCReadCounts: Estimation of cell-level SNVs from scRNA-seq data 93%
Similar papers in this journal
- Prediction of single-cell chromatin compartments from single-cell chromosome structures by MaxComp 95%
- Scuphr: A probabilistic framework for cell lineage tree reconstruction 95%
- A wavelet-based approach generates quantitative, scale-free and hierarchical descriptions of 3D genome structures and new biological insights 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.