deSAMBA: fast and accurate classification of metagenomics long reads with sparse approximate matches
Li, G.; Liu, B.; Wang, Y.
Show abstract
SummaryLong read sequencing technologies are promising to metagenomics studies. However, there is still lack of read classification tools to fast and accurately identify the taxonomies of noisy long reads, which is a bottleneck to the use of long read sequencing. Herein, we propose deSAMBA, a tailored long read classification approach that uses a novel sparse approximate match block (SAMB)-based pseudo alignment algorithm. Benchmarks on real datasets demonstrate that deSAMBA enables to simultaneously achieve fast speed and good classification yields, which outperforms state-of-the-art tools and has many potentials to cutting-edge metagenomics studies.\n\nAvailability and Implementationhttps://github.com/hitbc/deSAMBA.\n\nSupplementary information:
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- High-Quality Genomes of Nanopore Sequencing by Homologous Polishing 97%
- metaMIC: reference-free Misassembly Identification and Correction of de novo metagenomic assemblies 96%
- MetaBinner: a high-performance and stand-alone ensemble binning method to recover individual genomes from complex microbial communities 96%
Similar papers in this journal
- METAMVGL: a multi-view graph-based metagenomic contig binning algorithm by integrating assembly and paired-end graphs 96%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 95%
- SLR-superscaffolder: a de novo scaffolding tool for synthetic long reads using a top-to-bottom scheme 95%
Similar papers in this journal
- LRTK: A platform agnostic toolkit for linked-read analysis of both human genomes and metagenomes 95%
- Characterization and simulation of metagenomic nanopore sequencing data with Meta-NanoSim 94%
- Sequence Compression Benchmark (SCB) database - a comprehensive evaluation of reference-free compressors for FASTA-formatted sequences 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.