Do DNA and cytometric measures agree on genome sizes?
Gilbert, D.
Show abstract
Measurement of DNA contents of genomes is valuable for understanding genome biology, including assessments of genome assemblies, but it is not a trivial problem. Measuring contents of DNA shotgun reads is complicated by several factors: biological contents of genomes, laboratory methods, sequencing technology and computational processing. This compares and shares complications with cytometric measures of genome size and contents. There is an obvious discrepancy between cytometry and current long-read assemblies: assemblies average significantly below cytometric sizes. Measures of population changes within a species will control some of these complications. This report examines five species population sets with published cytometric and DNA data sets: Arabidopsis thaliana and A. arenosa, Arctic plants of Cochlearia genus, clonal populations of a rotifer Brachionus asplanchnoidis, and Zea mays corn plants. Results of this are clear, if complicated: DNA and cytometry do measure the same genome sizes, when done carefully with controls or adjustments for errors. Population changes in genome sizes are found by assembly-mapped measures of DNA, including environment or regional effects of latitude and altitude. Copy numbers of repeats, transposons and genes are changing. Kmer-based measures of DNA generally fail to match cytometry, miss population changes, and are opaque to understanding measurement errors. Oxford Nanopore technology produces the least biased DNA for measurement, with recent ONT.R10 Simplex data a match to cytometric sizes for corn, tomato plants, zebra fish and zebra finch bird. Assemblies of these species DNA average 12% below measured sizes, incomplete for duplicated content. New assembly of this ONT Simplex DNA reaches the size measured by cytometry and Gnodes, in 4 of 5 species. rRNA gene duplications measure one aspect of this discrepancy: the genome assemblies examined all are missing many rRNA genes. New experiments that measure both cytometry and DNA, controlling error factors, are warranted to clarify these results and suggest improvements for genome projects.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A haplotype-complete chromosome-level assembly of octoploid Urochloa humidicola cv. Tully reveals multiple genomic compositions and evolutionary histories in the species 93%
- A haplotype-resolved reference genome for Eucalyptus grandis 93%
- Occurrence of Aneuploidy Across the Range of Coast Redwood (Sequoia sempervirens) 93%
Similar papers in this journal
- <underline>Genome-wide Imputation Using the Practical Haplotype Graph in the Heterozygous Crop Cassava</underline> 92%
- The assembled and annotated genome of the pigeon louse Columbicola columbae, a model ectoparasite 92%
- Chromonomer: a tool set for repairing and enhancing assembled genomes through integration of genetic maps and conserved synteny 92%
Similar papers in this journal
- An exploration of assembly strategies and quality metrics on the accuracy of the Knightia excelsa (rewarewa) genome. 95%
- Chromosome-level hybrid de novo genome assemblies as an attainable option for non-model organisms 94%
- A second unveiling: haplotig masking of the eastern oyster genome improves population-level inference 94%
Similar papers in this journal
- Repetitive DNA content in the maize genome is uncoupled from population stratification at SNP loci 94%
- RecView: an interactive R application for viewing and locating recombination positions using pedigree data 92%
- High-fidelity (repeat) consensus sequences from short reads using combined read clustering and assembly 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.