Software choice and depth of sequence coverage can impact plastid genome assembly - A case study in the narrow endemic Calligonum bakuense
Giorgashvili, E.; Reichel, K.; Caswara, C.; Kerimov, V.; Borsch, T.; Gruenstaeudl, M.
Show abstract
Most plastid genome sequences are assembled from short-read whole-genome sequencing data, yet the impact that sequence coverage and the choice of assembly software can have on the accuracy of the resulting assemblies is poorly understood. In this study, we test the impact of both factors on plastid genome assembly in the threatened and rare endemic shrub Calligonum bakuense, which forms a distinct lineage in the genus Calligonum. We aim to characterize the differences across plastid genome assemblies generated by different assembly software tools and levels of sequence coverage and to determine if these differences are large enough to affect the phylogenetic position inferred for C. bakuense. Four assembly software tools (FastPlast, GetOrganelle, IOGA, and NOVOPlasty) and three levels of sequence coverage (original depth, 2,000x, and 500x) are compared in our analyses. The resulting assemblies are evaluated with regard to reproducibility, contig number, gene complement, inverted repeat length, and computation time; the impact of sequence differences on phylogenetic tree inference is also assessed. Our results show that software choice can have a considerable impact on the accuracy and reproducibility of plastid genome assembly and that GetOrganelle produced the most consistent assemblies for C. bakuense. Moreover, we found that a cap in sequence coverage can reduce both the sequence variability across assembly contigs and computation time. While no evidence was found that the sequence variability across assemblies was large enough to affect the phylogenetic position inferred for C. bakuense, differences among the assemblies may influence genotype recognition at the population level.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromosome-scale reference genome of Pectocarya recurvata, a species with one of the smallest genome sizes in Boraginaceae 96%
- PhyloHerb: A phylogenomic pipeline for processing genome skimming data for plants 96%
- HybPhaser: a workflow for the detection and phasing of hybrids in target capture datasets 95%
Similar papers in this journal
- Jack of all trades: genome assembly of Wild Jack and comparative genomics of Artocarpus 95%
- Chromosome-level genome assemblies for two quinoa inbred lines from northern and southern highlands of Altiplano where quinoa originated 95%
- Refining dual RNA-seq mapping: sequential and combined approaches in host-parasite plant dynamics 95%
Similar papers in this journal
- Genomic comparison of non-photosynthetic plants from the family Balanophoraceae with their photosynthetic relatives. 96%
- Comparative plastomics of Amaryllidaceae: Inverted repeat expansion and the degradation of the ndh genes in Strumaria truncata Jacq. 96%
- Plastid genomics of Nicotiana (Solanaceae): insights into molecular evolution, positive selection and the origin of the maternal genome of Aztec tobacco (Nicotiana rustica) 96%
Similar papers in this journal
- The Genome of Chenopodium ficifolium: Developing Genetic Resources and a Diploid Model System for Allotetraploid Quinoa 96%
- Conserving a threatened North American walnut: a chromosome-scale reference genome for butternut (Juglans cinerea) 96%
- Genomic diversity and evolution in the Hawaiian Islands endemic Kokia (Malvaceae) 96%
Similar papers in this journal
- Comparative Genomics of Six Juglans Species Reveals Patterns of Disease-associated Gene Family Contractions. 96%
- Genome diversity and phylogeny of the section Alatae of genus Lemna (Lemnaceae), comprising the presumed species Lemna aequinoctialis, Le. perpusilla and Le. aoukikusa 95%
- A revised view on the evolution of glutamine synthetase isoenzymes in plants 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.