RefKA: A fast and efficient long-read genome assembly approach for large and complex genomes
Yuan, Y.; Bayer, P. E.; Anderson, R.; Lee, H.; Chan, C.-K. K.; Zhao, R.; Batley, J.; Edwards, D.
Show abstract
Recent advances in long-read sequencing have the potential to produce more complete genome assemblies using sequence reads which can span repetitive regions. However, overlap based assembly methods routinely used for this data require significant computing time and resources. Here, we have developed RefKA, a reference-based approach for long read genome assembly. This approach relies on breaking up a closely related reference genome into bins, aligning k-mers unique to each bin with PacBio reads, and then assembling each bin in parallel followed by a final bin-stitching step. During benchmarking, we assembled the wheat Chinese Spring (CS) genome using publicly available PacBio reads in parallel in 168 wall hours on a 250 CPU system. The maximum RAM used was 300 Gb and the computing time was 42,000 CPU hours. The approach opens applications for the assembly of other large and complex genomes with much-reduced computing requirements. The RefKA pipeline is available at https://github.com/AppliedBioinformatics/RefKA
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A view of the pan-genome of domesticated cowpea (Vigna unguiculata Walp.) 94%
- A sorghum Practical Haplotype Graph facilitates genome-wide imputation and cost-effective genomic prediction 93%
- A second generation capture panel for cost-effective sequencing of genome regulatory regions in wheat and relatives 93%
Similar papers in this journal
- Sorghum pan-genome explores the functional utility to accelerate the genetic gain 94%
- Novel design of imputation-enabled SNP arrays for breeding and research applications supporting multi-species hybridisation 94%
- A Partially Phase-Separated Genome Sequence Assembly of the Vitis Rootstock 'Börner' (Vitis riparia x Vitis cinerea) and its Exploitation for Marker Development and Targeted Mapping 93%
Similar papers in this journal
- Genomic patterns of introgression in interspecific populations created by crossing wheat with its wild relative 94%
- Genomic region associated with pod color variation in pea (Pisum sativum) 94%
- De novo whole-genome assembly of Chrysanthemum makinoi, a key wild ancestor to hexaploid Chrysanthemum 93%
Similar papers in this journal
- High-quality chromosome-scale assembly of the walnut (Juglans regia L) reference genome 93%
- The draft nuclear genome assembly of Eucalyptus pauciflora: new approaches to comparing de novo assemblies 92%
- A graph clustering algorithm for detection and genotyping of structural variants from long reads 91%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.