Back

Scalable neighbour search and alignment with uvaia

de Oliveira Martins, L.; Mather, A. E.; Page, A. J.

2023-03-04 bioinformatics
10.1101/2023.01.31.526458 bioRxiv
Show abstract

Despite millions of SARS-CoV-2 genomes being sequenced and shared globally, manipulating such data sets is still challenging, especially selecting sequences for focused phylogenetic analysis. We present a novel method, uvaia, which is based on partial and exact sequence similarity for quickly extracting database sequences similar to query sequences of interest. Many SARS-CoV-2 phylogenetic analyses rely on very low numbers of ambiguous sites as a measure of quality since ambiguous sites do not contribute to single nucleotide polymorphism (SNP) differences, which uvaia alleviates by using measures of sequence similarity that consider partially ambiguous sites. Such fine-grained definition of similarity allows not only for better phylogenetic analyses, but also for improved classification and biogeographical inferences. Uvaia works natively with compressed files, can use multiple cores and efficiently utilises memory, being able to analyse large data sets on a standard desktop.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.