Centrifuge+: improving metagenomic analysis upon Centrifuge
Liu, J.; Ma, R.; Ren, Y.; Guo, H.
Show abstract
SummaryAccurate abundance estimation of species is essential for metagenomic analysis. Although many methods have been developed for classification of metagenomic data and abundance estimation of species, the abundance estimation of species remains challenging due to the ambiguous reads that align equally well to more than one genome. Here, we present Centrifuge+, which introduces unique mapping rate to describe the influence of similarities among species in the reference database when analyzing ambiguous reads. In contrast to the popular Centrifuge, Centrifuge+ improved the accuracy of abundance estimation on simulated reads from 4278 complete prokaryotic genomes. Availability and implementationThe source code is available at https://github.com/mNGSmethods/Centrifugep. Contacth.guo@foxmail.com or jlsljf0101@126.com Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- METAMVGL: a multi-view graph-based metagenomic contig binning algorithm by integrating assembly and paired-end graphs 96%
- TrieDedup: A fast trie-based deduplication algorithm to handle ambiguous bases in high-throughput sequencing 96%
- Persistent Memory as an Effective Alternative to Random Access Memory in Metagenome Assembly 95%
Similar papers in this journal
- StrainFLAIR: Strain-level profiling of metagenomic samples using variation graphs 95%
- Automated evaluation of multiple sequence alignment methods to handle third generation sequencing errors 93%
- DnoisE: Distance denoising by Entropy. An open-source parallelizable alternative for denoising sequence datasets 93%
Similar papers in this journal
Similar papers in this journal
- Real-time resolution of short-read assembly graph using ONT long reads 94%
- Deep6mA: a deep learning framework for exploring similar patterns in DNA N6-methyladenine sites across different species 94%
- An assembly-free method of phylogeny reconstruction using short-read sequences from pooled samples without barcodes 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.