Plastid Genome Assembly Using Long-read Data (ptGAUL)
Zhou, W.; Armijos, C.; Lee, C.; Lu, R.; Wang, J.; Ruhlman, T.; Jansen, R.; Jones, A.; Jones, C.
Show abstract
Although plastid genome (plastome) structure is highly conserved across most seed plants, investigations during the past two decades revealed several disparately related lineages that experienced substantial rearrangements. Most plastomes contain a large, inverted repeat and two single-copy regions and few dispersed repeats, however the plastomes of some taxa harbor long repeat sequences (>300 bp). These long repeats make it difficult to assemble complete plastomes using short-read data leading to misassemblies and consensus sequences that have spurious rearrangements. Single-molecule, long-read sequencing has the potential to overcome these challenges, yet there is no consensus on the most effective method for accurately assembling plastomes using long-read data. We generated a pipeline, plastid Genome Assembly Using Long-read data (ptGAUL), to address the problem of plastome assembly using long-read data from Oxford Nanopore Technologies (ONT) or Pacific Biosciences platforms. We demonstrated the efficacy of the ptGAUL pipeline using 16 published long-read datasets. We showed that ptGAUL produces accurate and unbiased assemblies. Additionally, we employed ptGAUL to assemble four new Juncus (Juncaceae) plastomes using ONT long reads. Our results revealed many long repeats and rearrangements in Juncus plastomes compared with basal lineages of Poales.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Jack of all trades: genome assembly of Wild Jack and comparative genomics of Artocarpus 96%
- Chromosome-level genome assemblies for two quinoa inbred lines from northern and southern highlands of Altiplano where quinoa originated 96%
- The genomic impact of mycoheterotrophy: targeted gene losses but extensive expression reprogramming. 95%
Similar papers in this journal
- Chromosome-scale reference genome of Pectocarya recurvata, a species with one of the smallest genome sizes in Boraginaceae 98%
- PhyloHerb: A phylogenomic pipeline for processing genome skimming data for plants 96%
- HybPhaser: a workflow for the detection and phasing of hybrids in target capture datasets 95%
Similar papers in this journal
- Evidence for plastome loss in the holoparasitic Mystropetalaceae 96%
- Impact of parasitic lifestyle and different types of centromere organization on chromosome and genome evolution in the plant genus Cuscuta 96%
- Diverging repeatomes in holoparasitic Hydnoraceae uncover a playground of genome evolution 96%
Similar papers in this journal
- Comparative Genomics of Six Juglans Species Reveals Patterns of Disease-associated Gene Family Contractions. 96%
- Whole Genome Assembly and Annotation of Northern Wild Rice, Zizania palustris L., Supports a Whole Genome Duplication in the Zizania Genus 96%
- Genome and transcriptome architecture of allopolyploid okra (Abelmoschus esculentus) 95%
Similar papers in this journal
- Intraspecific genome size variation in Rorippa indica reveals a tropical adaptation by genomic enlargement 93%
- Comparative Analysis of Small Secreted Peptide Signaling during Defense Response: Insights from Vascular and Non-Vascular Plants 93%
- Two independent loss-of-function mutations in anthocyanidin synthase homeologous genes make sweet basil all green 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.