Measure of major contents in animal and plant genomes, using Gnodes, finds under assemblies of model plant, Daphnia, fire ant and others.
Gilbert, D.
Show abstract
Significant discrepancies in genome sizes measured by cytometric methods versus DNA sequence estimates are frequent, including recent long-read DNA assemblies of plant and animal genomes. A new DNA sequence measure using a baseline of unique conserved genes, Gnodes, finds the larger cytometric measures are often accurate. DNA-informatic measures of size, as well as assembly methods, have errors in methodology that under-measure duplicated genome spans. Major contents of several model and discrepant genomes are assessed here, including human, corn, chicken, insects, crustaceans, and the model plant. Transposons dominate larger genomes, structural repeats are often a major portion of smaller ones. Gene coding sequences are found in similar amounts across the taxonomic spread. The largest contributors to size discrepancies are higher-order repeats, but duplicated coding sequences are a significant missed content, and transposons in some examined species. Informatics of measuring DNA and producing assemblies, including recent long-read telomere to telomere approaches, are subject to mistakes in operation and/or interpretation that are biased against repeats and duplications. Mistaken aspects include alignment methods that are inaccurate for high-copy duplicated spans; misclassification of true repetitive sequence as heterozygosity and artifact; software default settings that exclude high-copy DNA; and overly conservative data processing that reduces duplicated genomic spans. Re-assemblies with balanced methods recover the missing portions of problem genomes including model plant, water fleas and fire ant.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- An exploration of assembly strategies and quality metrics on the accuracy of the Knightia excelsa (rewarewa) genome. 96%
- Accuracy of de novo assembly of DNA sequences from double-digest libraries varies substantially among software 95%
- Chromosome-level hybrid de novo genome assemblies as an attainable option for non-model organisms 95%
Similar papers in this journal
Similar papers in this journal
- Long-Read Sequencing of the Zebrafish Genome Reorganizes Genomic Architecture 94%
- High-fidelity (repeat) consensus sequences from short reads using combined read clustering and assembly 94%
- Telomere-to-telomere assembly of the genome of an individual Oikopleura dioica from Okinawa using Nanopore-based sequencing 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.