Mitochondrial sequences or Numts - By-catch differs between sequencing methods
Becher, H.; Nichols, R. A.
Show abstract
Nuclear inserts derived from mitochondrial DNA (Numts) encode valuable information. Being mostly non-functional, and accumulating mutations more slowly than mitochondrial sequence, they act like molecular fossils - they preserve information on the ancestral sequences of the mitochondrial DNA. In addition, changes to the Numt sequence since their insertion into the nuclear genome carry information about the nuclear phylogeny. These attributes cannot be reliably exploited if Numt sequence is confused with the mitochondrial genome (mtDNA). The analysis of mtDNA would be similarly compromised by any confusion, for example producing misleading results in DNA barcoding that used mtDNA sequence. We propose a method to distinguish Numts from mtDNA, without the need for comprehensive assembly of the nuclear genome or the physical separation of organelles and nuclei. It exploits the different biases of long and short-read sequencing. We find that short-read data yield mainly mtDNA sequences, whereas long-read sequencing strongly enriches for Numt sequences. We demonstrate the method using genome-skimming (coverage < 1x) data obtained on Illumina short-read and PacBio long-read technology from DNA extracted from six grasshopper individuals. The mitochondrial genome sequences were assembled from the short-read data despite the presence of Numts. The PacBio data contained a much higher proportion of Numt reads (over 16-fold), making us caution against the use of long-read methods for studies using mitochondrial loci. We obtained two estimates of the genomic proportion of Numts. Finally, we introduce \"tangle plots\", a way of visualising Numt structural rearrangements and comparing them between samples.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- MitoFinder: efficient automated large-scale extraction of mitogenomic data in target enrichment phylogenomics 95%
- A snakemake toolkit for the batch assembly, annotation, and phylogenetic analysis of mitochondrial genomes and ribosomal genes from genome skims of museum collections. 95%
- Chromosome-level hybrid de novo genome assemblies as an attainable option for non-model organisms 95%
Similar papers in this journal
- Nanopore genome skimming with Illumina polishing yields highly accurate mitogenome sequences: a case study of Niphargus amphipods 96%
- High quality genome assembly of the brown hare (Lepus europaeus) with chromosome-level scaffolding 95%
- Chromosome-level reference genome assembly for the mountain hare (Lepus timidus) 95%
Similar papers in this journal
- Repeat-rich regions cause false positive detection of NUMTs - a case study in amphibians using an improved cane toad reference genome 96%
- Giants among Cnidaria: large nuclear genomes and rearranged mitochondrial genomes in siphonophores 96%
- Evolutionary Insights from the Mitochondrial Genome of Oikopleura dioica: Sequencing Challenges, RNA Editing, Gene Transfers to the Nucleus, and tRNA Loss 96%
Similar papers in this journal
- NUMT PARSER: automated identification and removal of nuclear mitochondrial pseudogenes (numts) for accurate mitochondrial genome reconstruction in Panthera 96%
- Comparison of whole-genome assemblies of European river lamprey (Lampetra fluviatilis) and brook lamprey (Lampetra planeri) 95%
- High-quality genome assembly of the endemic threatened White-bellied Sholakili Sholicola albiventris (Muscicapidae: Blanford, 1868) from the Shola Sky Islands, India. 95%
Similar papers in this journal
- Evolutionary history of an Alpine archaeognath(Machilis pallida) - insights from different variant types 96%
- A Prelude to Conservation Genomics: First Chromosome-Level Genome Assembly of a Flying Squirrel (Pteromyini: Pteromys volans) 94%
- Covering the bases: population genomic structure of Lemna minor and the cryptic species L. japonica in Switzerland 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.