MADRe: Strain-Level Metagenomic Classification Through Assembly-Driven Database Reduction
Lipovac, J.; Sikic, M.; Vicedomini, R.; Krizanovic, K.
Show abstract
AbstractStrain-level metagenomic classification is essential for understanding microbial diversity and functional potential, but remains challenging, par- ticularly in the absence of prior knowledge about the composition of the sample. In this paper we present MADRe, a modular and scalable pipeline for long-read strain-level metagenomic classification, enhanced with Metagenome Assembly-Driven Database Reduction. MADRe com- bines long-read metagenome assembly, contig-to-reference mapping reas- signment based on an expectation-maximization algorithm for database reduction, and probabilistic read mapping reassignment to achieve sensi- tive and precise classification. We extensively evaluated MADRe on sim- ulated datasets, mock communities, and a real anaerobic digester sludge metagenome, demonstrating that it consistently outperforms existing tools by achieving higher precision with reduced false positives. MADRes de- sign allows users to apply either the database reduction or read classi- fication step individually. Using only the read classification step shows results on par with other tested tools. MADRe is open source and pub- licly available at https://github.com/lbcb-sci/MADRe.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Benchmarking Alignment Strategies for Hi-C Reads in Metagenomic Hi-C Data 96%
- LEMMIv2: Benchmarking Framework for Metagenomic and 16S Amplicon Profilers with a Catalogue of Evaluated Tools. 96%
- HiCBin: Binning metagenomic contigs and recovering metagenome-assembled genomes using Hi-C contact maps 96%
Similar papers in this journal
- Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets 97%
- Functional Analysis of Metagenomes by Likelihood Inference (FAMLI) Successfully Compensates for Multi-Mapping Short Reads from Metagenomic Samples 95%
- Sketching and sampling approaches for fast and accurate long read classification 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.