Back

Recombinogenic G-quadruplexes in the Newtonian DNA Sequence Space

Kuryavyi, V. V.

2026-07-03 genomics
10.64898/2026.06.30.735570 bioRxiv
Show abstract

Abstract The universe of possible nucleotide sequences expands combinatorially with sequence length, vastly exceeding the fraction sampled by real genomes. Yet genomic sequences exhibit reproducible compositional symmetries and recurrent structural motifs, indicating that biological sequence space is shaped by strong organizing constraints. Here, we introduce an explicit framework for constructing and visualizing the complete sequence universe using the Newtonian polynomial for a four-letter alphabet, and for identifying biologically relevant subsets through the application of fundamental filters. Three filters of biological relevance are formulated: (i) the constraint that DNA predominantly exists as an antiparallel-stranded double helix, (ii) the second Chargaff parity rule, which enforces approximate strand symmetry in single-stranded sequence composition, and (iii) genome shadows, reflecting the imprint of concerted sequence changes. Successive application of these filters dramatically reduces the accessible sequence space and reveals distinct symmetry classes. Among these, mirror-symmetric sequences occupy a privileged position because they are invariant under strand reversal and therefore compatible with both antiparallel and parallel strand orientations. This dual compatibility enables such sequences to bridge otherwise disjoint structural subspaces of DNA. G-rich members of this class are shown to have a strong propensity to form G-quadruplex architectures that incorporate parallel-stranded domains while remaining compatible with duplex DNA. We propose that this structural versatility provides a mechanistic basis for the recurrent association of G-rich mirror-symmetric sequences with recombination hotspots and genome rearrangements. Together, these results establish a symmetry-based framework for understanding how combinatorial sequence space is filtered into biologically functional DNA motifs.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

1
Nucleic Acids Research
1281 papers in training set
Top 0.3%
30.5%
2
Proceedings of the National Academy of Sciences
2444 papers in training set
Top 2%
12.4%
3
Nature Communications
5641 papers in training set
Top 17%
10.8%
50% of probability mass above
4
Nature Structural & Molecular Biology
18 papers in training set
Top 0.1%
4.3%
5
PLOS Computational Biology
1863 papers in training set
Top 8%
4.2%
6
Cell
431 papers in training set
Top 3%
3.2%
7
eLife
5828 papers in training set
Top 42%
2.3%
8
Genome Research
468 papers in training set
Top 3%
2.3%
9
NAR Genomics and Bioinformatics
242 papers in training set
Top 2%
2.1%
10
Genome Biology
637 papers in training set
Top 5%
2.1%
11
Molecular Biology and Evolution
542 papers in training set
Top 3%
2.1%
12
Nature
645 papers in training set
Top 6%
1.9%
13
GENETICS
483 papers in training set
Top 3%
1.7%
14
Cell Reports
1498 papers in training set
Top 20%
1.7%
15
Cell Systems
201 papers in training set
Top 4%
1.1%
16
Scientific Reports
3612 papers in training set
Top 67%
1.1%
17
Science
477 papers in training set
Top 7%
1.1%
18
Nature Biotechnology
172 papers in training set
Top 4%
1.0%
19
Journal of The Royal Society Interface
235 papers in training set
Top 4%
1.0%
20
PNAS Nexus
159 papers in training set
Top 4%
0.8%
21
Science Advances
1243 papers in training set
Top 31%
0.8%
22
Nature Genetics
286 papers in training set
Top 5%
0.8%
23
Molecular Cell
350 papers in training set
Top 6%
0.6%
24
Molecular Systems Biology
162 papers in training set
Top 4%
0.6%
25
Genome Biology and Evolution
338 papers in training set
Top 4%
0.6%