aaKomp: Alignment-free amino acid k-mer matching for genome completeness assessment at scale
Wong, J.; Coombe, L.; Warren, R. L.; Birol, I.
Show abstract
In de novo sequencing projects, genome assembly optimization requires evaluating a number of candidate assemblies to identify optimal tool parameters. Yet, current completeness assessment tools like BUSCO and compleasm require 10-80 minutes per evaluation for gigabase-scale genomes, transforming what should be rapid iteration into time-intensive processes. These tools rely on alignment-based approaches and fixed ortholog databases, limiting their scalability across the tree of life. We present aaKomp, a scalable alignment-free tool that leverages amino acid k-mer matching and multi-index Bloom filters for rapid genome completeness assessment. Unlike current utilities, aaKomp supports user-defined reference databases, enabling customized assessments for any organism. In benchmarking against state-of-the-art tools using simulated T2T-CHM13 datasets, aaKomp achieved 68-fold faster execution and 15-fold lower memory consumption while maintaining accuracy. Testing on 94 Human Pangenome Reference Consortium assemblies and a European Eel assembly, aaKomp maintained one-minute runtimes (1.2 {+/-} 0.35 min) and low memory usage (<13.64 GB). aaKomps scoring system provides nuanced estimates rather than threshold-based classifications, offering increased resolution for tracking incremental improvements during iterative workflows. aaKomps speed, memory efficiency, and flexible database generation makes it well-suited for modern and biodiverse projects requiring the evaluation of hundreds of assemblies.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- BRAKER3: Fully Automated Genome Annotation Using RNA-Seq and Protein Evidence with GeneMark-ETP, AUGUSTUS and TSEBRA 96%
- HiCanu: accurate assembly of segmental duplications, satellites, and allelic variants from high-fidelity long reads 96%
- Geometric deep learning framework for de novo genome assembly 95%
Similar papers in this journal
Similar papers in this journal
- iLoci: Robust evaluation of genome content and organization for provisional and mature genome assemblies 96%
- PyOrthoANI, PyFastANI, and Pyskani: a suite of Python libraries for computation of average nucleotide identity 95%
- MoGAAAP: A modular Snakemake workflow for automated genome assembly and annotation with quality assessment 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.