Estimation of substitution and indel rates via k-mer statistics
HERA, M. R.; Medvedev, P.; Koslicki, D.; Blanca, A.
Show abstract
Methods utilizing k-mers are widely used in bioinformatics, yet our understanding of their statistical properties under realistic mutation models remains incomplete. Previously, substitution-only mutation models have been considered to derive precise expectations and variances for mutated k-mers and intervals of mutated and non-mutated sequences. In this work, we consider a mutation model that incorporates insertions and deletions in addition to single-nucleotide substitutions. Within this framework, we derive closed-form k-mer-based estimators for the three fundamental mutation parameters: substitution, deletion rate, and insertion rates. We provide theoretical guarantees in the form of concentration inequalities, ensuring accuracy of our estimators under reasonable model assumptions. Empirical evaluations on simulated evolution of genomic sequences confirm our theoretical findings, demonstrating that accounting for insertions and deletions signals allows for accurate estimation of mutation rates and improves upon the results obtained by considering a substitution-only model. An implementation of estimating the mutation parameters from a pair of fasta files is available here: github.com/KoslickiLab/estimate_rates_using_mutation_model.git. The results presented in this manuscript can be reproduced using the code available here: github.com/KoslickiLab/est_rates_experiments.git. 2012 ACM Subject ClassificationApplied computing [->] Computational biology; Theory of computation [->] Theory and algorithms for application domains; Mathematics of computing [->] Probabilistic inference problems
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Debiasing FracMinHash and deriving confidence intervals for mutation rates across a wide range of evolutionary distances 98%
- Sequence aligners can guarantee accuracy in almost O(m log n) time: a rigorous average-case analysis of the seed-chain-extend heuristic 97%
- Strobemers: an alternative to k-mers for sequence comparison 96%
Similar papers in this journal
- The statistics of k-mers from a sequence undergoing a simple mutation process without spurious matches 98%
- Enabling inference for context-dependent models of mutation by bounding the propagation of dependency 96%
- Studying the history of tumor evolution from single-cell sequencing data by exploring the space of binary matrices 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.