Fairy: fast approximate coverage for multi-samplemetagenomic binning
Shaw, J.; Yu, Y. W.
Show abstract
BackgroundMetagenomic binning, the clustering of assembled contigs that belong to the same genome, is a crucial step for recovering metagenomeassembled genomes (MAGs). Contigs are linked by exploiting consistent read coverage patterns across a genome. Using coverage from multiple samples leads to higher-quality MAGs; however, standard pipelines require all-to-all read alignments for multiple samples to compute coverage, becoming a key computational bottleneck. ResultsWe present fairy (https://github.com/bluenote-1577/fairy), an approximate coverage calculation method for metagenomic binning. Fairy is a fast k-mer-based alignment-free method. For multi-sample binning, fairy can be > 250x faster than read alignment and accurate enough for binning. Fairy is compatible with several existing binners on host and non-host-associated datasets. Using MetaBAT2, fairy recovers 98.5% of MAGs with > 50% completeness and < 5% incompleteness relative to alignment with BWA. Notably, multi-sample binning with fairy is always better than single-sample binning using BWA (> 1.5x more > 50% complete MAGs on average) while still being faster. For a public sediment metagenome project, we demonstrate that multisample binning recovers higher quality Asgard archaea MAGs than single-sample binning and that fairys results are indistinguishable from read alignment. ConclusionsFairy is a new tool for approximately and quickly calculating multi-sample coverage for binning, resolving a longstanding computational bottleneck for metagenomics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Sketching and sampling approaches for fast and accurate long read classification 96%
- Evaluation of taxonomic classification and profiling methods for long-read shotgun metagenomic sequencing datasets 95%
- HapSolo: An optimization approach for removing secondary haplotigs during diploid genome assembly and scaffolding. 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.