The number of k-mer matches between two DNA sequences as a function of k and applications to estimate phylogenetic distances
Röhling, S.; Linne, A.; Schellhorn, J.; Hosseini, M.; Dencker, T.; Morgenstern, B.
Show abstract
We study the number Nk of (spaced) word matches between pairs of evolutionarily related DNA sequences depending on the word length or pattern weight k, respectively. We show that, under the Jukes-Cantor model, the number of substitutions per site that occurred since two sequences evolved from their last common ancestor, can be esti-mated from the slope of a certain function of Nk. Based on these considerations, we implemented a software program for alignment-free sequence comparison called Slope-SpaM. Test runs on simulated sequence data show that Slope-SpaM can estimate phylogenetic dis-tances with high accuracy for up to around 0.5 substitutions per po-sitions. The statistical stability of our results is improved if spaced words are used instead of contiguous k-mers. Unlike previous methods that are based on the number of (spaced) word matches, our approach can deal with sequences that share only local homologies.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- The statistics of k-mers from a sequence undergoing a simple mutation process without spurious matches 97%
- Enabling inference for context-dependent models of mutation by bounding the propagation of dependency 97%
- Determining significant correlation between pairs of extant characters in a small parsimony framework 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.