Back

Wengan: Efficient and high quality hybrid de novo assembly of human genomes

Di Genova, A.; Buena-Atienza, E.; Ossowski, S.; Sagot, M.-F.

2019-11-25 bioinformatics
10.1101/840447 bioRxiv
Show abstract

The continuous improvement of long-read sequencing technologies along with the development of ad-doc algorithms has launched a new de novo assembly era that promises high-quality genomes. However, it has proven difficult to use only long reads to generate accurate genome assemblies of large, repeat-rich human genomes. To date, most of the human genomes assembled from long error-prone reads add accurate short reads to further polish the consensus quality. Here, we report the development of a novel algorithm for hybrid assembly, WO_SCPCAPENGANC_SCPCAP, and the de novo assembly of four human genomes using a combination of sequencing data generated on ONT PromethION, PacBio Sequel, Illumina and MGI technology. WO_SCPCAPENGANC_SCPCAP implements efficient algorithms that exploit the sequence information of short and long reads to tackle assembly contiguity as well as consensus quality. The resulting genome assemblies have high contiguity (contig NG50:16.67-62.06 Mb), few assembly errors (contig NGA50:10.9-45.91 Mb), good consensus quality (QV:27.79-33.61), and high gene completeness (BO_SCPCAPUSCOC_SCPCAP complete: 94.6-95.1%), while consuming low computational resources (CPU hours:153-1027). In particular, the WO_SCPCAPENGANC_SCPCAP assembly of the haploid CHM13 sample achieved a contig NG50 of 62.06 Mb (NGA50:45.91 Mb), which surpasses the contiguity of the current human reference genome (GRCh38 contig NG50:57.88 Mb). Providing highest quality at low computational cost, WO_SCPCAPENGANC_SCPCAP is an important step towards the democratization of the de novo assembly of human genomes. The WO_SCPCAPENGANC_SCPCAP assembler is available at https://github.com/adigenova/wengan

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.