Back

XTree enables memory-efficient, accurate short and long sequence alignment to millions of genomes across the tree of life

Al-Ghalith, G. A.; Ryon, K. A.; Henriksen, J. R.; Danko, D. C.; Farthing, B.; Marengo, M.; Church, G.; Peixoto, R.; Patel, C. J.; Knights, D.; Tierney, B. T.

2025-12-29 bioinformatics
10.64898/2025.12.22.696015 bioRxiv
Show abstract

XTree is a k-mer-based aligner enabling rapid, memory-efficient alignment of sequencing reads to whole-genome reference databases with up to millions of genomes. Here, we detail XTrees performance on short and long read sequencing data and demonstrate its high accuracy across diverse bacterial, viral, and eukaryotic genomes. Benchmarking demonstrates superior and/or comparable precision and recall over existing tools, with more scalable indexing and efficient memory mapping. We additionally provide pre-indexed databases, including (1) the Genome Taxonomy Database (versions r214-r226), (2) representative GenBank fungi and protozoan genomes and (3) the Pan-Viral-Compendium, a bespoke data resource spanning 6.6 million, quality-controlled, viral genomes.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.