Benchmarking of bioinformatics tools for the hybrid de novo assembly of human whole-genome sequencing data
Munoz-Barrera, A.; Rubio-Rodriguez, L. A.; Jaspez, D.; Corrales, A.; Marcelino-Rodriguez, I.; Lorenzo-Salazar, J. M.; Gonzalez-Montelongo, R.; Flores, C.
Show abstract
Accurate and complete de novo assembled genomes sustain variant identification and catalyze the discovery of new genomic features and biological functions. However, accurate and precise de novo assemblies of large and complex genomes remains a challenging task. Long-read sequencing data alone or in hybrid mode combined with more accurate short-read sequences facilitate the de novo assembly of genomes. A number of software exists for de novo genome assembly from long-read data although specific performance comparisons to assembly human genomes are lacking. Here we benchmarked 11 different pipelines including four long-read only assemblers and three hybrid assemblers, combined with four polishing schemes for de novo genome assembly of a human reference material sequenced with Oxford Nanopore Technologies and Illumina. In addition, the best performing choice was validated in a non-reference routine laboratory sample. Software performance was evaluated by assessing the quality of the assemblies with QUAST, BUSCO, and Merqury metrics, and the computational costs associated with each of the pipelines were also assessed. We found that Flye was superior to all other assemblers, especially when relying on Ratatosk error-corrected long-reads. Polishing improved the accuracy and continuity of the assemblies and the combination of two rounds of Racon and Pilon achieved the best results. The assembly of the non-reference sample showed comparable assembly metrics as those of the reference material. Based on the results, a complete optimal analysis pipeline for the assembly, polishing, and contig curation developed on Nextflow is provided to enable efficient parallelization and built-in dependency management to further advance in the generation of high-quality and chromosome-level human assemblies.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- High-fidelity (repeat) consensus sequences from short reads using combined read clustering and assembly 94%
- MicroPIPE: An end-to-end solution for high-quality complete bacterial genome construction 94%
- Fine-Tuning GBS Data with Comparison of Reference and Mock Genome Approaches for Advancing Genomic Selection in Less Studied Farmed Species 94%
Similar papers in this journal
- Choice of assemblers has a critical impact on de novo assembly of SARS-CoV-2 genome and characterizing variants 95%
- From contigs towards chromosomes: automatic Improvement of Long Read Assemblies (ILRA) 95%
- SatXplor - A comprehensive pipeline for satellite DNA analyses in complex genome assemblies 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.