Back

Reference-free resolution of long-read metagenomic data

Khachatryan, L.; Anvar, S. Y.; Vossen, R. H. A. M.; Laros, J. F. J.

2019-10-21 bioinformatics
10.1101/811760 bioRxiv
Show abstract

BackgroundRead binning is a key step in proper and accurate analysis of metagenomics data. Typically, this is performed by comparing metagenomics reads to known microbial sequences. However, microbial communities usually contain mixtures of hundreds to thousands of unknown bacteria. This restricts the accuracy and completeness of alignment-based approaches. The possibility of reference-free deconvolution of environmental sequencing data could benefit the field of metagenomics, contributing to the estimation of metagenome complexity, improving the metagenome assembly, and enabling the investigation of new bacterial species that are not visible using standard laboratory or alignment-based bioinformatics techniques.\n\nResultsHere, we apply an alignment-free method that leverages on k-mer frequencies to classify reads within a single long read metagenomic dataset. In addition to a series of simulated metagenomic datasets, we generated sequencing data from a bioreactor microbiome using the PacBio RSII single-molecule real-time sequencing platform. We show that distances obtained after the comparison of k-mer profiles can reveal relationships between reads within a single metagenome, leading to a clustering per species.\n\nConclusionsIn this study, we demonstrated the possibility to detect substructures within a single metagenome operating only with the information derived from the sequencing reads. The obtained results are highly important as they establish a principle that might potentially expand the toolkit for the detection and investigation of previously unknow microorganisms.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.