Can Random Walking on a Hi-C Contact Matrix Lead to Data Quality Improvement? An Assessment
Lin, S.; Liu, Y.
Show abstract
Hi-C and single cell Hi-C (scHi-C) data are now routinely generated for studying an array of biological questions of interest, including whole genome chromatin organization to gain a better understanding of the chromosome three-dimensional hierarchical structure: compartments, Topologically Associated Domains (TADs), and long-range interactions. Due to concerns about data quality, especially for scHi-C because of its sparsity, data quality improvement is seen as a necessary step before performing analyses to answer biological questions. As such, methods have been developed accordingly, among them is a set of methods that are "random walk"-based, including random walk with a limited number of steps (RWS) and random walk with restart (RWR). Nevertheless, there is little justification for the use of such methods, nor quantification of their performance success. Taking correct identification of TADs as the end point, in this paper, we describe the characteristics of random-walk-based approaches and carry out empirical investigation for identifying TADs before and after random walks. Due to the lack of practical guidelines for choosing tuning parameters necessary for performing random walks, it is difficult to know how many steps of random walk for RWS or how small a restart probability for RWR should one choose to achieve good performance. Even in the unrealistic scenario when one has the hindsight of using the optimal parameter values, little improvement in downstream studies by first performing random walk was observed. This conclusion was based on extensive analytical analyses, simulation study, and real data applications. Therefore, the current study provides a cautionary note to researchers who may consider using random-walk-based approaches prior to downstream analyses.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
- Simultaneous smoothing and detection of topological units of genome organization from sparse chromatin contact count matrices with matrix factorization 96%
- Accuracy, Robustness and Scalability of Dimensionality Reduction Methods for Single Cell RNAseq Analysis 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.