Two dimensional sequence alignment shows that replication slippage may generate a significant proportion of all transversion substitutions.
Erives, A. J.
Show abstract
Homologous sequences diverge in length via insertions and deletions (indels). Consequently, evolutionary genetic analyses routinely use methods to produce gapped alignment (GA). In GA, artificial null characters (gaps) are inserted into sequences so that nucleotide characters may be placed into homological correspondence within an alignment column. However, this approach sacrifices the homological correspondence of nucleotides diverging via tandem repeats (TRs). To address this deficit, we generalize GA with micro-paralogical gapped alignment (MPGA). While GA operates under a strict two-state homology model of one-to-one and one-to-none (i.e. one-to-gap) relationships, MPGA adds one-to-many, many-to-many, and many-to-none relationships. This expanded, multi-state homology model is motivated by DNA replication slippage (RS). RS produces short tandem repeats, constituting interrelated micro-paralogous sequences. Together, RS and TR-associated instability have a synergistic effect in the production of indels, which generate the need for gap insertions. MPGA reduces the computational cost of determining optimal gap insertions by reducing the number of gaps required by two-dimensional (2D) representations of sequence. A 2D representation of one sequence is achieved when tandem repeats are contracted into the same columns (dimension one) by occupying multiple rows (dimension two), an internal micro-paralogical dimension. To demonstrate the benefits and challenges of 2D representation, we develop a program called LINEUP and identify a pervasive fractal dimension in evolving sequences. We then demonstrate how LINEUP-generated 2D representations provide improved measures of substitution rates and transition-to-transversion ratios. Altogether, these results showcase significant new perspectives on basic mutational and evolutionary processes when multi-state homology models are adopted.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Machine learning techniques for classifying the mutagenic origins of point mutations 94%
- Natural variation in codon bias and mRNA folding strength interact synergistically to modify protein expression in Saccharomyces cerevisiae 94%
- UV damage induces production of mitochondrial DNA fragments with specific length profiles 93%
Similar papers in this journal
- Evolutionarily conserved non-protein-coding regions in the chicken genome harbor functionally important variation 94%
- A Genetic Screen for Suppressors of Cryptic 5' Splicing in C. elegans Reveals Roles for KIN17 and PRCC in Maintaining Both 5' and 3' Splice Site Identity 94%
- Sequence determinants and evolution of constitutive and alternative splicing in yeast species 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.