Anniemap: Vector Search for Viral Short Read Alignment
van Zyl, D. J.; Tegally, H.; Baxter, C.; The INFORM Africa research study group, ; de Oliveira, T.; Xavier, J. S.; Dunaiski, M.
Show abstract
Background: The process of aligning sequencing reads to a reference genome is a foundational step in genomic analysis, underpinning tasks from variant detection to pathogen surveillance. In viral genomics, however, this problem becomes substantially more challenging: viral sequences are often present at low abundance within host-dominated samples and can differ markedly from available references due to rapid mutation and population heterogeneity. These characteristics reduce the effectiveness of conventional seed-and-extend aligners, which typically rely on long exact or near-exact matches to anchor alignments. Even modest sequence divergence or sequencing errors can disrupt such seeds, particularly for short reads, leading to missed alignments. The central challenge in this setting is maintaining robust alignment under high divergence without sacrificing efficiency. Results: We introduce Anniemap, a vector search based approach to viral short-read sequence alignment. Anniemap represents reads and reference sequences as binary vectors and performs approximate nearest-neighbour search using Facebook AI Similarity Search (FAISS) to efficiently identify candidate mappings. Anniemap was compared with the well-established alignment tools Bowtie2 and BWA-MEM2 across a diverse set of viral genomes and read lengths using both simulated and real sequencing data. Anniemap achieved higher sensitivity and throughput in almost all evaluated scenarios, with the most substantial improvements in sensitivity observed for highly divergent genomes, such as Hepatitis C virus (HCV) and Human Immunodeficiency Virus (HIV). Conclusions; By measuring vector similarity rather than relying on long exact seed matches, Anniemap provides greater robustness to sequencing errors and genomic mutations. This property is particularly advantageous for viral genomes, where substantial sequence divergence is common. Further work is required to efficiently extend vector-based search for read alignment beyond viral genomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Influenza Classification from Short Reads with VAPOR Facilitates Robust Mapping Pipelines and Zoonotic Strain Detection for Routine Surveillance Applications 95%
- nf-core/viralmetagenome: A Novel Pipeline for Untargeted Viral Genome Reconstruction 94%
- Capturing variation in metagenomic assembly graphs with MetaCortex 93%
Similar papers in this journal
- BINSEQ: A Family of High-Performance Binary Formats for Nucleotide Sequences 92%
- ConNIS and labeling instability: new statistical methods for improving the detection of essential genes in TraDIS libraries 92%
- Demonstrating the utility of flexible sequence queries against indexed short reads with FlexTyper 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.