ProGenFixer: an ultra-fast and accurate tool for correcting prokaryotic genome sequences using a mapping-free algorithm
Song, L.
Show abstract
Although prokaryotic genomes are simpler, making variation analysis more straightforward, an efficient and user-friendly tool for their rapid sequence correction is lacking. We present ProGenFixer, an ultra-fast, mapping-free tool specifically designed to efficiently identify and correct errors in prokaryotic genomes using next-generation sequencing (NGS) data. ProGenFixer compares k-mer profiles between the reference genome and reads to pinpoint discrepancies, employs a local assembly-based algorithm to estimate corrected sequences, and automatically implements corrections. Benchmarking demonstrates ProGenFixers speed and accuracy advantages. It is over >5x faster than the swiftest existing tool tested while maintaining higher accuracy. Compared to slower, highly accurate tools, ProGenFixer shows comparable accuracy but is >17x faster and demonstrates enhanced performance on long indels (>10 bp). Implemented in C, ProGenFixer is a standalone, user-friendly tool offering a comprehensive solution for prokaryotic genome correction. The software is freely available at https://github.com/Scilence2022/ProGenFixer. Additionally, a companion web-based platform offering a ready-to-use interface can be accessed at https://progenfixer.biodesign.ac.cn.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- EvANI benchmarking workflow for evolutionary distance estimation 95%
- A Computational Toolset for Rapid Identification of SARS-CoV-2, other Viruses, and Microorganisms from Sequencing Data 95%
- Hound: A novel tool for automated mapping of genotype to phenotype in bacterial genomes assembled de novo 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.