Hybrid Clustering of Long and Short-read for Improved Metagenome Assembly
Lu, Y.; Shi, L.; Van Goethem, M. W.; Sevim, V.; Mascagni, M.; Deng, L.; Wang, Z.
Show abstract
Next-generation sequencing has enabled metagenomics, the study of the genomes of microorganisms sampled directly from the environment without cultivation. We previously developed a proof-of-concept, scalable metagenome clustering algorithm based on Apache Spark to cluster sequence reads according to their species of origin. To overcome its under-clustering problem on short-read sequences, in this study we developed a new, two-step Label Propagation Algorithm (LPA) that first forms clusters of long reads and then recruits short reads to these clusters. Compared to alternative label propagation strategies, this hybrid clustering algorithm (hybrid-LPA) yields significantly larger read clusters without compromising cluster purity. We show that adding an extra clustering step before assembly leads to improved metagenome assemblies, predicting more complete genomes or gene clusters from a synthetic metagenome dataset and a real-world metagenome dataset, respectively. These results suggest that hybrid-LPA is a good alternative to current metagenome assembly practice by providing benefits in both scalability and accuracy on large metagenome datasets. Availability and implementationhttps://bitbucket.org/zhong_wang/hybridlpa/src/master/. Contactzhongwang@lbl.gov
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Persistent Memory as an Effective Alternative to Random Access Memory in Metagenome Assembly 97%
- SQMtools: automated processing and visual analysis of 'omics data with R and anvi'o 96%
- METAMVGL: a multi-view graph-based metagenomic contig binning algorithm by integrating assembly and paired-end graphs 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.