Towards the extended barcode concept: Generating DNA reference data through genome skimming of danish plants
Chua, P. Y. S.; Leerhoi, F.; Langkjaer, E. M. R.; Margaryan, A.; Noer, C. L.; Richter, S. R.; Restrup, M. E.; Bruun, H. H.; Hartvig, I.; Coissac, E.; Boessenkool, S.; Alsos, I. G.; Bohmann, K.
Show abstract
BackgroundRecently, there has been a push towards the extended barcode concept of utilising chloroplast genomes (cpGenome) and nuclear ribosomal DNA (nrDNA) sequences for molecular identification of plants instead of the standard barcode regions. These extended barcodes has a wide range of applications, including biodiversity monitoring and assessment, primer design, and evolutionary studies. However, these extended barcodes are not well represented in global reference databases. To fill this gap, we generated cpGenomes and nrDNA reference data from genome skims of 184 plant species collected in Denmark. We further explored the application of our generated reference data for molecular identifications of plants in an environmental DNA metagenomics study. ResultsWe assembled partial cpGenomes for 82.1% of sequenced species and full or partial nrDNA sequences for 83.7% of species. We added all assemblies to GenBank, of which chloroplast reference data from 101 species and nuclear reference data from 6 species were not previously represented. On average, we recovered 45 genes per species. The rate of recovery of standard barcodes was higher for nuclear barcodes (>89%) than chloroplast barcodes (< 60%). Extracted DNA yield did not affect assembly outcome, whereas high GC content did so negatively. For the in silico simulation of metagenomic reads, taxonomic assignments using the reference data generated had better species resolution (94.9%) as compared to GenBank (18.1%) without any identification errors. ConclusionsGenome skimming generates reference data of both standard barcodes and other loci, contributing to the global DNA reference database for plants.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromosome-level de novo genome assembly of Telopea speciosissima (New South Wales waratah) using long-reads, linked-reads and Hi-C 95%
- A multispecies amplicon sequencing approach for genetic diversity assessment in grassland plant species 94%
- An exploration of assembly strategies and quality metrics on the accuracy of the Knightia excelsa (rewarewa) genome. 94%
Similar papers in this journal
- Draft genome assemblies using sequencing reads from Oxford Nanopore Technology and Illumina platforms for four species of North American killifish from the Fundulus genus 95%
- The draft nuclear genome assembly of Eucalyptus pauciflora: new approaches to comparing de novo assemblies 95%
- Long-reads assembly of the Brassica napus reference genome, Darmor-bzh 95%
Similar papers in this journal
- Can we use it? On the utility of de novo and reference-based assembly of Nanopore data for plant plastome sequencing 98%
- Genetic variations associated with adaptation processes in Acrocomia palms: A comparative study across the Neotropic for future crop improvement 95%
- A 3K Axiom(R) SNP array from a transcriptome-wide SNP resource sheds new light on the genetic diversity and structure of the iconic subtropical conifer tree Araucaria angustifolia (Bert.) Kuntze 95%
Similar papers in this journal
- Draft genome of the aquatic moss Fontinalis antipyretica (Fontinalaceae, Bryophyta) 96%
- A reference assembly for the legume cover crop, hairy vetch (Vicia villosa) 96%
- Chromosome-scale assembly of the highly heterozygous genome of red clover (Trifolium pratense L.), an allogamous forage crop species 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.