CoCoBin: Graph-Based Metagenomic Binning via Composition-Coverage Separation
Potiwara, K. -; Wichadakul, D. -
Show abstract
MotivationMetagenomic binning is a critical step in metagenomic analysis, aiming to cluster contigs from the same genome into coherent groups. In contemporary workflows, most binning tools begin with the assembly of shotgun metagenomic sequencing data. The assembled contigs are then grouped into bins representing individual microbial genomes or species, typically using taxonomy-independent methods. Although several methods exist, metagenomic binning remains a challenging yet mandatory task, particularly in the context of complex and highly diverse microbial communities. ResultsWe propose CoCoBin, a novel metagenomic binning tool explicitly designed for the effective binning of metagenomic contigs. In this study, we introduced an innovative approach for calculating contig similarity by separating composition and coverage information. The method begins by (1) assigning contigs into a cluster based on length ranges, (2) calculating contig similarity based on composition features (e.g., k-mer frequencies), and (3) calculating contig difference based on coverage features. These similarity measures are then integrated to construct a graph, where nodes represent contigs and edges represent the similarities between them. Finally, the Louvain algorithm is applied to the graph to cluster closely related contigs. CoCoBin was compared against several state-of-the-art binning tools: BusyBee Web, CONCOCT, MaxBin 2.0, MetaBAT 2, and MetaDecoder on nine simulated datasets, five mock community datasets, and one real dataset. The AMBER tool used to evaluate the binning results across all datasets shows that CoCoBin achieved the best performance regarding the number of bins identified, followed by its performance on the F1 score. AvailabilityThe source code of CoCoBin is available at https://github.com/cucpbioinfo/CoCoBin Contactduangdao.w@chula.ac.th Supplementary informationSupplementary data are available at Bioinformatics online.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- SQMtools: automated processing and visual analysis of 'omics data with R and anvi'o 97%
- METAMVGL: a multi-view graph-based metagenomic contig binning algorithm by integrating assembly and paired-end graphs 96%
- Persistent Memory as an Effective Alternative to Random Access Memory in Metagenome Assembly 96%
Similar papers in this journal
Similar papers in this journal
- Deciphering active prophages from metagenomes 95%
- MetaPhage: an automated pipeline for analyzing, annotating, and classifying bacteriophages in metagenomics sequencing data. 95%
- Decomposing a San Francisco Estuary microbiome using long read metagenomics reveals species and species- and strain-level dominance from picoeukaryotes to viruses 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.