Minipoa: A minimizer-based method for fast and memory-efficient partial order alignment
Liu, H.; Zhang, P.; Wei, Y.; Tian, Q.; Zhai, Y.; Zou, Q.; Niu, M.
Show abstract
Partial order alignment (POA) has emerged as a fundamental component in long-read error correction, assembly and pangenomics. However, conventional POA algorithms are limited by high time and memory requirements, making them inefficient for large-scale datasets. Here, we present minipoa, a fast and memory-efficient POA tool that incorporates seed-chain-align heuristics, adaptive or static banding strategies, and single-instruction multiple-data optimizations. Minipoa achieves up to a 5-fold speedup over abPOA, reduces memory usage by up to 16-fold, and improves correction accuracy, while maintaining strong performance on both PacBio and ONT simulated datasets, and can be readily integrated into existing long-read error correction and assembly workflows. In multiple sequence alignment datasets, minipoa demonstrates superior computational efficiency and alignment accuracy compared with all other tested tools, achieving Total Column scores up to 2.5-fold higher than MAFFT in low-similarity scenarios. Moreover, minipoa enables multiple sequence alignment of megabase-long genomes and million-sequence datasets, demonstrated by 342 Mycobacterium tuberculosis sequences and one million SARS-CoV-2 sequences respectively. Collectively, minipoa is well positioned to become a cornerstone in the era of large-scale pangenomics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- EvANI benchmarking workflow for evolutionary distance estimation 96%
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 95%
- Clair3-Trio: high-performance Nanopore long-read variant calling in family trios with Trio-to-Trio deep neural networks 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.