Polymer-derived distance penalties improve chromatin interaction predictions from single-cell data across crop genomes
Schlegel, L.; Gomez-Cano, F.; Marand, A. P.; Johannes, F.
Show abstract
Scalable proxies for 3D genome contacts - such as single-cell co-accessibility and deep learning predictions - have emerged as powerful alternatives to chromatin capture-based methods, but predictions systematically overestimate long-range interactions. Here we show how to correct this bias using distance-based penalty functions informed by Gaussian mixture modeling and polymer-physics scaling. Using Hi-C datasets from maize, rice, and soybean, we derive tissue-specific and global consensus penalties parameterized by multi-regime power-law exponents. Applying these corrections to scATAC-seq co-accessibility scores improves their distance profiles in concordance with Hi-C and reduces long-range false positives by an average of 73% with tissue-specific penalties and 66% with the global consensus. We provide open-source code and fitted parameters to support adoption in maize, rice, and soybean.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Developing a general AI model for integrating diverse genomic modalities and comprehensive genomic knowledge 95%
- Deciphering the 3D genome organization across species from Hi-C data 95%
- Simultaneous Profiling of Chromatin Accessibility and DNA Methylation in Complete Plant Genomes Using Long-Read Sequencing 95%
Similar papers in this journal
Similar papers in this journal
- Geometrically encoded positioning of introns, intergenic segments, and exons in the human genome 96%
- Unveiling multi-scale architectural features in single-cell Hi-C data using scCAFE 94%
- Hierarchical prediction and perturbation of chromatin organization reveal how loop domains mediate higher-order architectures 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.