Scalable neighbour search and alignment with uvaia
de Oliveira Martins, L.; Mather, A. E.; Page, A. J.
Show abstract
Despite millions of SARS-CoV-2 genomes being sequenced and shared globally, manipulating such data sets is still challenging, especially selecting sequences for focused phylogenetic analysis. We present a novel method, uvaia, which is based on partial and exact sequence similarity for quickly extracting database sequences similar to query sequences of interest. Many SARS-CoV-2 phylogenetic analyses rely on very low numbers of ambiguous sites as a measure of quality since ambiguous sites do not contribute to single nucleotide polymorphism (SNP) differences, which uvaia alleviates by using measures of sequence similarity that consider partially ambiguous sites. Such fine-grained definition of similarity allows not only for better phylogenetic analyses, but also for improved classification and biogeographical inferences. Uvaia works natively with compressed files, can use multiple cores and efficiently utilises memory, being able to analyse large data sets on a standard desktop.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rapid screening and detection of inter-type viral recombinants using phylo-k-mers 96%
- PoSeiDon: a Nextflow pipeline for the detection of evolutionary recombination events and positive selection 96%
- MLDSP-GUI: An alignment-free standalone tool with an interactive graphical user interface for DNA sequence comparison and analysis 96%
Similar papers in this journal
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 96%
- Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors 95%
- Rephine.r: a pipeline for correcting gene calls and clusters to improve phage pangenomes and phylogenies 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.