MAGE: Strain Level Profiling of MetagenomeSamples
Walia, V.; V.G, S.; Srinivasan, R.; Sivadasan, N.
Show abstract
Metagenomic profiling from sequencing data aims to disentangle a microbial sample at lower ranks of taxonomy, such as species and strains. Deep taxonomic profiling involving accurate estimation of strain level abundances aids in precise quantification of the microbial composition, which plays a crucial role in various downstream analyses. Existing tools primarily focus on strain/subspecies identification and limit abundance estimation to the species level. Abundance quantification of the identified strains is challenging and remains largely unaddressed by the existing approaches. We propose a novel algorithm MAGE (Microbial Abundance GaugE), for accurately identifying constituent strains and quantifying strain level relative abundances. For accurate profiling, MAGE uses read mapping information and performs a novel local searchbased profiling guided by a constrained optimization based on maximum likelihood estimation. Unlike the existing approaches that often rely on strain-specific markers and homology information for deep profiling, MAGE works solely with read mapping information, which is the set of target strains from the reference collection for each mapped read. As part of MAGE, we provide an alignment-free and kmer-based read mapper that uses a compact and comprehensive index constructed using FM-index and R-index. We use a variety of evaluation metrics for validating abundances estimation quality. We performed several experiments using a variety of datasets, and MAGE exhibited superior performance compared to the existing tools on a wide range of performance metrics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- AmpliDiff: An Optimized Amplicon Sequencing Approach to Estimating Lineage Abundances in Viral Metagenomes 97%
- Keeping up with the genomes: efficient learning of our increasing knowledge of the tree of life 96%
- Read-SpaM: assembly-free and alignment-free comparison of bacterial genomes with low sequencing coverage 96%
Similar papers in this journal
- MetaMLP: A fast word embedding based classifier to profile target gene databases in metagenomic samples 95%
- Metabolic pathway prediction using non-negative matrix factorization with improved precision 94%
- An Efficient, Scalable and Exact Representation of High-Dimensional Color Information Enabled via de Bruijn Graph Search 94%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.