Back

Quadrupia: Derivation of G-quadruplexes for organismal genomes across the tree of life

Chantzi, N.; Nayak, A.; Baltoumas, F. A.; Aplakidou, E.; Liew, S. W.; Galuh, J. E.; Patsakis, M.; Moeckel, C.; Mouratidis, I.; Sazed, S. A.; Guiblet, W.; Montgomery, A.; Karmiris-Obratanski, P.; Guliang, W.; Zaravinos, A.; Vasquez, K. M.; Kwok, C. K.; Pavlopoulos, G.; Georgakopoulos-Soares, I.

2024-07-11 genomics
10.1101/2024.07.09.602008 bioRxiv
Show abstract

G-quadruplex DNA structures exhibit a profound influence on essential biological processes, including transcription, replication, telomere maintenance, and genomic stability. These structures have demonstrably shaped organismal evolution. However, a comprehensive, organism-wide G-quadruplex map encompassing the diversity of life has remained elusive. Here, we introduce Quadrupia, the most extensive and well-characterized G-quadruplex database to date, facilitating the exploration of G-quadruplex structures across the evolutionary spectrum. Quadrupia has identified G-quadruplex sequences in 108,449 reference genomes, with a total of 140,181,277 G-quadruplexes. The database also hosts a collection of 319,784 G-quadruplex clusters of 20 or more members, annotated by taxonomic distributions, multiple sequence alignments, profile Hidden Markov Models and cross-references to G-quadruplex 3D structures. Examination of G-quadruplexes across functional genomic elements in different taxa indicates preferential orientation and positioning, with significant differences between individual taxonomic groups. For example, we find that G-quadruplexes in bacteria with a single replication origin display profound preference for the leading orientation. Finally, we experimentally validate the most frequently observed G-quadruplexes using CD-spectroscopy, UV melting, and fluorescent-based approaches. Quadrupia is publicly available through https://www.pavlopoulos-lab.org/quadrupia.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.