Investigating the performance of deep learning methods for Hi-C resolution improvement
Murtaza, G.; Jain, A.; Hughes, M.; Varatharajan, T.; Singh, R.
Show abstract
MotivationHi-C is a widely used technique to study the 3D organization of the genome. Due to its high sequencing cost, most of the generated datasets are of coarse resolution, which makes it impractical to study finer chromatin features such as Topologically Associating Domains (TADs) and chromatin loops. Multiple deep-learning-based methods have recently been proposed to increase the resolution of these data sets by imputing Hi-C reads (typically called upscaling). However, the existing works evaluate these methods on either synthetically downsampled or a small subset of experimentally generated sparse Hi-C datasets, making it hard to establish their generalizability in the real-world use case. We present our framework - Hi-CY - that compares existing Hi-C resolution upscaling methods on seven experimentally generated low-resolution Hi-C datasets belonging to various levels of read sparsities originating from three cell lines on a comprehensive set of evaluation metrics. Hi-CY also includes four downstream analysis tasks, such as TAD and chromatin loops recall, to provide a thorough report on the generalizability of these methods. ResultsWe observe that existing deep-learning methods fail to generalize to experimentally generated sparse Hi-C datasets showing a performance reduction of up to 57 %. As a potential solution, we find that retraining deep-learning based methods with experimentally generated Hi-C datasets improves performance by up to 31%. More importantly, Hi-CY shows that even with retraining, the existing deep-learning based methods struggle to recover biological features such as chromatin loops and TADs when provided with sparse Hi-C datasets. Our study, through Hi-CY framework, highlights the need for rigorous evaluation in future. We identify specific avenues for improvements in the current deep learning-based Hi-C upscaling methods, including but not limited to using experimentally generated datasets for training. Availabilityhttps://github.com/rsinghlab/Hi-CY Author SummaryWe evaluate deep learning-based Hi-C upscaling methods with our framework Hi-CY using seven datasets originating from three cell lines evaluated using three correlation metrics, four Hi-C similarity metrics, and four downstream analysis tasks, including TAD and chromatin loop recovery. We identify a distributional shift between Hi-C contact matrices generated from downsampled and experimentally generated sparse Hi-C datasets. We use Hi-CY to establish that the existing methods trained with downsampled Hi-C datasets tend to perform significantly worse on experimentally generated Hi-C datasets. We explore potential strategies to alleviate the drop in performance such as retraining models with experimentally generated datasets. Our results suggest that retraining improves performance up to 31 % on five sparse GM12878 datsets but provides marginal improvement in cross cell-type setting. Moreover, we observe that regardless of the training scheme, all deep-learning based methods struggle to recover biological features such as TADs and chromatin loops when provided with very sparse experimentally generated datasets as inputs.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Graph Convolutional Networks for Epigenetic State Prediction Using Both Sequence and 3D Genome Data 97%
- seqgra: Principled Selection of Neural Network Architectures for Genomics Prediction Tasks 96%
- Consensus Label Propagation with Graph Convolutional Networks for Single-Cell RNA Sequencing Cell Type Annotation 96%
Similar papers in this journal
- Tuning parameters of dimensionality reduction methods for single-cell RNA-seq analysis 97%
- scDesign2: a transparent simulator that generates high-fidelity single-cell gene expression count data with gene correlations captured 96%
- Enhancement of network architecture alignment in comparative single-cell studies 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.