SkewDB: A comprehensive database of GC and a 10 other skews for over 28,000 chromosomes and plasmids
Hubert, B.
Show abstract
GC skew denotes the relative excess of G nucleotides over C nucleotides on the leading versus the lagging replication strand of eubacteria. While the effect is small, typically around 2.5%, it is robust and pervasive. GC skew and the analogous TA skew are a localized deviation from Chargaffs second parity rule, which states that G and C, and T and A occur with (mostly) equal frequency even within a strand. Most bacteria also show the analogous TA skew. Different phyla show different kinds of skew and differing relations between TA and GC skew. This article introduces an open access database (https://skewdb.org) of GC and 10 other skews for over 28,000 chromosomes and plasmids. Further details like codon bias, strand bias, strand lengths and taxonomic data are also included. The SkewDB database can be used to generate or verify hypotheses. Since the origins of both the second parity rule, as well as GC skew itself, are not yet satisfactorily explained, such a database may enhance our understanding of microbial DNA.
Matching journals
The top 11 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Extraction of near-complete genomes from metagenomic samples: a new service in PATRIC 94%
- GAMBIT (Genomic Approximation Method for Bacterial Identification and Tracking): A methodology to rapidly leverage whole genome sequencing of bacterial isolates for clinical identification 94%
- DeLUCS: Deep Learning for Unsupervised Clustering of DNA Sequences 93%
Similar papers in this journal
Similar papers in this journal
- Visualizing and quantifying structural diversity around mobile resistance genes 93%
- Whokaryote: distinguishing eukaryotic and prokaryotic contigs in metagenomes based on gene structure 92%
- Platon: identification and characterization of bacterial plasmid contigs in short-read draft assembliesexploiting protein-sequence-based replicon distribution scores 92%
Similar papers in this journal
- NanoSpring: reference-free lossless compression of nanopore sequencing reads using an approximate assembly approach 92%
- Detecting SARS-CoV-2 lineages and mutational load in municipal wastewater; a use-case in the metropolitan area of Thessaloniki, Greece 92%
- Identifying the essential genes of Mycobacterium avium subsp. hominissuis with Tn-Seq using a rank-based filter procedure. 91%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.