An open-sourced bioinformatic pipeline for the processing of Next-Generation Sequencing derived nucleotide reads: Identification and authentication of ancient metagenomic DNA.
Collin, T. C.; Drosou, K.; O'Riordan, J. D.; Meshveliani, T.; Pinhasi, R.; Feeney, R. N. M.
Show abstract
Bioinformatic pipelines optimised for the processing and assessment of metagenomic ancient DNA (aDNA) are needed for studies that do not make use of high yielding DNA capture techniques. These bioinformatic pipelines are traditionally optimised for broad aDNA purposes, are contingent on selection biases and are associated with high costs. Here we present a bioinformatic pipeline optimised for the identification and assessment of ancient metagenomic DNA without the use of expensive DNA capture techniques. Our pipeline actively conserves aDNA reads, allowing the application of a bioinformatic approach by identifying the shortest reads possible for analysis (22-28bp). The time required for processing is drastically reduced through the use of a 10% segmented non-redundant sequence file (229 hours to 53). Processing speed is improved through the optimisation of BLAST parameters (53 hours to 48). Additionally, the use of multi-alignment authentication in the identification of taxa increases overall confidence of metagenomic results. DNA yields are further increased through the use of an optimal MAPQ setting (MAPQ 25) and the optimisation of the duplicate removal process using multiple sequence identifiers (a 4.35-6.88% better retention). Moreover, characteristic aDNA damage patterns are used to bioinformatically assess ancient vs. modern DNA origin throughout pipeline development. Of additional value, this pipeline uses open-source technologies, which increases its accessibility to the scientific community.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Detection of introduced and resident marine species using environmental DNA metabarcoding of sediment and water 94%
- Monitoring fish communities through environmental DNA metabarcoding in the fish pass system of the second largest hydropower plant in the world 94%
- Gene-flow from steppe individuals into Cucuteni-Trypillia associated populations indicates long-standing contacts and gradual admixture 94%
Similar papers in this journal
- Highly accurate long-read HiFi sequencing data for five complex genomes 95%
- Marine picoplankton metagenomes from eleven vertical profiles obtained by the Malaspina Expedition in the tropical and subtropical oceans 95%
- Characterizing organisms from three domains of life with universal primers from throughout the global ocean 94%
Similar papers in this journal
- Transcriptomic stability or lability explains sensitivity to climate stressors in coralline algae 93%
- Long-insert sequence capture detects high copy numbers in a defence-related beta-glucosidase gene Betaglu-1 with large variations in white spruce but not Norway spruce 93%
- A highly contiguous genome assembly of Brassica nigra (BB) and revised nomenclature for the pseudochromosomes 91%
Similar papers in this journal
- Shark and ray genome size estimation: methodological optimization for inclusive and controllable biodiversity genomics 94%
- Nanopore long reads enable the first complete genome assembly of a Malaysian Vibrio parahaemolyticus isolate bearing the pVa plasmid associated with acute hepatopancreatic necrosis disease 92%
- NAD: Noise-augmented direct sequencing of target nucleic acids by augmenting with noise and selective sampling 90%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.