StrainPro: a highly accurate Metagenomic strain-level profiling tool
Lin, H.-N.; Lin, Y.-L.; Hsu, W.-L.
Show abstract
Characterizing the taxonomic diversity of a microbial community is very important to understand the roles of microorganisms. Next generation sequencing (NGS) provides great potential for investigation of a microbial community and leads to Metagenomic studies. NGS generates DNA fragment sequences directly from microorganism samples, and it requires analysis tools to identify microbial species (or taxonomic composition) and estimate their relative abundance in the studied community. However, only a few tools could achieve strain-level identification and most tools estimate the microbial abundances simply according to the read counts. An evaluation study on metagenomic analysis tools concludes that the predicted abundance differed significantly from the true abundance. In this study, we present StrainPro, a novel metagenomic analysis tool which is highly accurate both at characterizing microorganisms at strain-level and estimating their relative abundances. A unique feature of StrainPro is it identifies representative sequence segments from reference genomes. We generate three simulated datasets using known strain sequences and another three simulated datasets using unknown strain sequences. We compare the performance of StrainPro with seven existing tools. The results show that StrainPro not only identifies metagenomes with high precision and recall, but it is also highly robust even when the metagenomes are not included in the reference database. Moreover, StrainPro estimates the relative abundance with high accuracy. We demonstrate that there is a strong positive linear relationship between observed and predicted abundances.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life 97%
- Comprehensive benchmarking of metagenomic classification tools for long-read sequencing data 96%
- METAMVGL: a multi-view graph-based metagenomic contig binning algorithm by integrating assembly and paired-end graphs 96%
Similar papers in this journal
- MetaMLP: A fast word embedding based classifier to profile target gene databases in metagenomic samples 96%
- HiCzin: Normalizing metagenomic Hi-C data and detecting spurious contacts using zero-inflated negative binomial regression 95%
- MetaCoAG: Binning Metagenomic Contigs via Composition, Coverage and Assembly Graphs 93%
Similar papers in this journal
- Variant Evolution Graph: Can We Infer How SARS-CoV-2 Variants are Evolving? 95%
- Extraction of near-complete genomes from metagenomic samples: a new service in PATRIC 95%
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.