Back

aaKomp: Alignment-free amino acid k-mer matching for genome completeness assessment at scale

Wong, J.; Coombe, L.; Warren, R. L.; Birol, I.

2026-03-22 bioinformatics
10.64898/2026.03.19.713078 bioRxiv
Show abstract

In de novo sequencing projects, genome assembly optimization requires evaluating a number of candidate assemblies to identify optimal tool parameters. Yet, current completeness assessment tools like BUSCO and compleasm require 10-80 minutes per evaluation for gigabase-scale genomes, transforming what should be rapid iteration into time-intensive processes. These tools rely on alignment-based approaches and fixed ortholog databases, limiting their scalability across the tree of life. We present aaKomp, a scalable alignment-free tool that leverages amino acid k-mer matching and multi-index Bloom filters for rapid genome completeness assessment. Unlike current utilities, aaKomp supports user-defined reference databases, enabling customized assessments for any organism. In benchmarking against state-of-the-art tools using simulated T2T-CHM13 datasets, aaKomp achieved 68-fold faster execution and 15-fold lower memory consumption while maintaining accuracy. Testing on 94 Human Pangenome Reference Consortium assemblies and a European Eel assembly, aaKomp maintained one-minute runtimes (1.2 {+/-} 0.35 min) and low memory usage (<13.64 GB). aaKomps scoring system provides nuanced estimates rather than threshold-based classifications, offering increased resolution for tracking incremental improvements during iterative workflows. aaKomps speed, memory efficiency, and flexible database generation makes it well-suited for modern and biodiverse projects requiring the evaluation of hundreds of assemblies.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.