Estimating chromosome sizes from karyotype images enables validation of de novo assemblies.
Ludwig, A.; Dibrov, A.; Myers, G.; Pippel, M.
Show abstract
Highly contiguous genome assemblies are essential for genomic research. Chromosome-scale assembly is feasible with the modern sequencing techniques in principle, but in practice, scaffolding errors frequently occur, leading to incorrect number and sizes of chromosomes. Relating the observed chromosome sizes from karyotype images to the generated assembly scaffolds offers a method for detecting these errors. Here, we present KICS, a semi-automated approach for estimating relative chromosome sizes from karyotype images and their subsequent comparison to the corresponding assembly scaffolds. The method relies on threshold-based image segmentation and uses the computed areas of the chromosome-related connected components as a proxy for the actual chromosome size. We demonstrate the validity and practicality of our approach by applying it to karyotype images of humans and various amphibians, birds, fish, insects, mammals, and plants. We found a strong linear relationship between pixel counts and the DNA content of chromosomes. Averaging estimates from eight human karyotype images, KICS predicts most of the chromosome sizes within an error margin of just 6 Mb. Our method provides additional means of validating genome assemblies at low costs. An interactive implementation of KICS is available at https://github.com/mpicbg-csbd/napari-kics.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- AlcoR: alignment-free simulation, mapping, and visualization of low-complexity regions in biological data 95%
- Smash++: an alignment-free and memory-efficient tool to find genomic rearrangements 95%
- ntsm: an alignment-free, ultra low coverage, sequencing technology agnostic, intraspecies sample comparison tool for sample swap detection 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.