Exhaustive benchmarking of de novo assembly methods for eukaryotic genomes
Southwood, D.; Rane, R. V.; Lee, S. F.; Oakeshott, J. G.; Ranganathan, S.
Show abstract
The assembly of reference-quality, chromosome-resolution genomes for both model and novel eukaryotic organisms is an increasingly achievable task for single research teams. However, the overwhelming abundance of sequencing technologies, assembly algorithms, and post-assembly processing tools currently available means that there is no clear consensus on a best-practice computational protocol for eukaryotic de novo genome assembly. Here, we provide a comprehensive benchmark of 28 state-of-the-art assembly and polishing packages, in various combinations, when assembling two eukaryotic genomes using both next-generation (Illumina HiSeq) and third-generation (Oxford Nanopore and PacBio CLR) sequencing data, at both controlled and open levels of sequencing coverage. Recommendations are made for the most effective tools for each sequencing technology and the best performing combinations of methods, evaluated against common assessment metrics such as contiguity, computational performance, gene completeness, and reference reconstruction, across both organisms and across sequencing coverage depth.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- MicroPIPE: An end-to-end solution for high-quality complete bacterial genome construction 96%
- Consistent ultra-long DNA sequencing with automated slow pipetting 94%
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 94%
Similar papers in this journal
- Chromosome-level quality scaffolding of brown algal genomes using InstaGRAAL, a proximity ligation-based scaffolder 96%
- Identifying and quantifying isoforms from accurate full-length transcriptome sequencing reads with Mandalorion 95%
- Benchmarking Alignment Strategies for Hi-C Reads in Metagenomic Hi-C Data 94%
Similar papers in this journal
- binny: an automated binning algorithm to recover high-quality genomes from complex metagenomic datasets 95%
- Choice of assemblers has a critical impact on de novo assembly of SARS-CoV-2 genome and characterizing variants 94%
- Comprehensive benchmarking of software for mapping whole genome bisulfite data: from read alignment to DNA methylation analysis 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.