Systematic assessment of sequencing depth requirements for Hi-C-derived metrics
Granovsky, A.; Polovnikov, K. E.
Show abstract
BackgroundHi-C experiments produce genome-wide chromatin contact maps from which structural features can be quantified across multiple genomic scales, ranging from megabase-scale compartments to kilobase-scale boundaries and chromatin loops. Although high-resolution analyses commonly rely on hundreds of millions of sequenced read pairs, the minimum sequencing depth required for different classes of Hi-C-derived features has not been systematically established. We therefore sought to determine these depth requirements systematically. ResultsWe performed progressive random subsampling of twelve Hi-C libraries representing multiple cell types and experimental protocols. Loop-density and loop-size inference from the P(s) log-derivative was evaluated across all libraries, whereas compartment, insulation, and boundary analyses were performed on a subset of nine libraries with comparable full-depth coverage. Rather than defining sufficient depth as an arbitrary fraction of the full-depth value, we introduced a biologically motivated criterion based on the variability between independent biological replicates. The required sequencing depth for each metric was defined as the point at which subsampling scatter first reached the variability observed between independent biological replicates, representing the accuracy that additional sequencing cannot improve upon. Using this criterion, loop-density estimation reached its biological floor at approximately 10 million read pairs, loop-size estimation at approximately 20 million, insulation scores and boundary detection at approximately 30 million, and compartment eigenvectors at 100-kilobase resolution at approximately 60 million read pairs. ConclusionsDifferent Hi-C-derived metrics require substantially different sequencing depths to achieve biologically meaningful accuracy. For all metrics, depth-induced variability fell below biological replicate variability well before full sequencing depth was reached. These thresholds provide practical guidance for experimental design and sequencing budget allocation, suggesting that, in many studies, increasing the number of biological replicates is likely to improve reproducibility more effectively than sequencing individual libraries to greater depth.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Characterizing the properties of bisulfite sequencing data: maximizing power and sensitivity to identify between-group differences in DNA methylation 91%
- Genome-wide Enhancer Maps Differ Significantly in Genomic Distribution, Evolution, and Function 90%
- BandHiC: a memory-efficient and user-friendly Python package for organizing and analyzing Hi-C matrices down to sub-kilobase resolution 90%
Similar papers in this journal
- Analysis of the structural variability of topologically associated domains as revealed by Hi-C 92%
- HIPPIE2: a method for fine-scale identification of physically interacting chromatin regions 92%
- Genome wide clustering on integrated chromatin states and Micro-C contacts reveals chromatin interaction signatures 92%
Similar papers in this journal
- Assessment of human diploid genome assembly with 10x Linked-Reads data 91%
- Identifying, understanding, and correcting technical biases on the sex chromosomes in next-generation sequencing data 90%
- epialleleR: an R/Bioconductor package for sensitive allele-specific methylation analysis in NGS data 89%
Similar papers in this journal
- preciseTAD: A transfer learning framework for 3D domain boundary prediction at base-pair resolution 93%
- Beyond Blacklists: A Critical Assessment of Exclusion Set Generation Strategies and Alternative Approaches 93%
- Chicdiff: a computational pipeline for detecting differential chromosomal interactions in Capture Hi-C data 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.