Reference-free resolution of long-read metagenomic data
Khachatryan, L.; Anvar, S. Y.; Vossen, R. H. A. M.; Laros, J. F. J.
Show abstract
BackgroundRead binning is a key step in proper and accurate analysis of metagenomics data. Typically, this is performed by comparing metagenomics reads to known microbial sequences. However, microbial communities usually contain mixtures of hundreds to thousands of unknown bacteria. This restricts the accuracy and completeness of alignment-based approaches. The possibility of reference-free deconvolution of environmental sequencing data could benefit the field of metagenomics, contributing to the estimation of metagenome complexity, improving the metagenome assembly, and enabling the investigation of new bacterial species that are not visible using standard laboratory or alignment-based bioinformatics techniques.\n\nResultsHere, we apply an alignment-free method that leverages on k-mer frequencies to classify reads within a single long read metagenomic dataset. In addition to a series of simulated metagenomic datasets, we generated sequencing data from a bioreactor microbiome using the PacBio RSII single-molecule real-time sequencing platform. We show that distances obtained after the comparison of k-mer profiles can reveal relationships between reads within a single metagenome, leading to a clustering per species.\n\nConclusionsIn this study, we demonstrated the possibility to detect substructures within a single metagenome operating only with the information derived from the sequencing reads. The obtained results are highly important as they establish a principle that might potentially expand the toolkit for the detection and investigation of previously unknow microorganisms.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Phylogeny analysis of whole protein-coding genes in metagenomic data detected an environmental gradient for the microbiota 95%
- Omnicrobe, an open-access database of microbial habitats and phenotypes using a comprehensive text mining and data fusion approach 95%
- Comparative evaluation of bioinformatic tools for virus-host prediction and their application to a highly diverse community in the Cuatro Cienegas Basin, Mexico 95%
Similar papers in this journal
- Assembly methods for nanopore-based metagenomic sequencing: a comparative study 97%
- Comparison of the effectiveness of different normalization methods for metagenomic cross-study phenotype prediction under heterogeneity 95%
- Interactive Analysis of Biosurfactants in Fruit-Waste Fermentation Samples using BioSurfDB and MEGAN 95%
Similar papers in this journal
- Dancing the Nanopore limbo - Nanopore metagenomics from small DNA quantities for bacterial genome reconstruction 96%
- Assessing the performance of different approaches for functional and taxonomic annotation of metagenomes 96%
- Metagenomic assemblies tend to break around antibiotic resistance genes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.