KOMB: Taxonomy-oblivious Characterization of Metagenome Dynamics via K-core Decomposition
Balaji, A.; Sapoval, N.; Elworth, R. A. L.; Segarra, S.; Treangen, T. J.
Show abstract
Characterizing metagenomic samples via kmer-based, database-dependent taxonomic classification methods has provided crucial insight into underlying host-associated microbiome dynamics. However, novel approaches are needed that are able to track microbial community dynamics within metagenomes to elucidate genome flux in response to perturbations and disease states. Here we describe KOMB, a novel approach for tracking homologous regions within microbiomes. KOMB utilizes K-core graph decomposition on metagenome assembly graphs to identify repetitive and homologous regions to varying degrees of resolution. K-core performs a hierarchical decomposition which partitions the graph into shells containing nodes having degree at least K, called K-shells, yielding O(V + E) complexity compared to exact betweenness centrality complexity of O(V E) found in prior related approaches. We show through rigorous validation on simulated, synthetic, and real metagenomic datasets that KOMB accurately recovers and profiles repetitive and homologous genomic regions across organisms in the sample. KOMB can also identify functionally-rich regions in Human Microbiome Project (HMP) datasets, and can be used to analyze longitudinal data and identify pivotal taxa in fecal microbiota transplantation (FMT) samples. In summary, KOMB represents a novel approach to microbiome characterization that can efficiently identify sequences of interest in metagenomes.
Matching journals
The top 6 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- CoCoPyE: feature engineering for learning and prediction of genome quality indices 96%
- Pangenome databases provide superior host removal and mycobacteria classification from clinical metagenomic data 96%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.