Digital Kennison: A bioinformatics pipeline for rapid mapping of sequences to the Drosophila melanogaster Y chromosome
Uno, F.; Carvalho, A. B.
Show abstract
The Drosophila melanogaster Y chromosome is currently known to contain 13 single-copy protein-coding genes, six of which are essential for male fertility, as well as several non-coding genes and abundant repetitive DNA. Localization of Y-linked sequences has traditionally relied on labor-intensive crosses using Kennisons translocation strains, which map Y-linked loci by generating flies deficient for each of the six Y-chromosome fertility regions (ks-1, ks-2, kl-1, kl-2, kl-3, and kl-5). Here we present Digital Kennison, a computational pipeline that recasts this classical mapping strategy as a sequence-based analysis. The pipeline queries eight genomic databases derived from Kennisons strains using BLAST and read coverage, assigning sequences to fertility regions with a calibrated confidence score. We benchmarked the method on 60 Y-linked sequences spanning all six regions, including single-copy protein-coding genes, Mst77Y family members, non-coding RNAs, and the centromere. Digital Kennison achieved 97% precision while resolving challenging cases, including boundary-spanning genes (PRY and Ppr-Y), fragmented Mst77Y copies, and FDY, which has a closely related autosomal paralog. Beyond validating known localizations, the pipeline localized the unmapped gene CG41561 to the kl-1region and reassigned the transcript CR40629-RC from the kl-2 region to kl-5. It also localized 7 of 16 recently transferred Y-linked sequences described by Tobler et al. (2017), including 4 with high confidence. Applied to 904 small R6 scaffolds, Digital Kennison assigned 75% to fertility regions, including five currently annotated as autosomal-pericentromeric. Digital Kennison reduces sequence localization from weeks of genetic crosses to minutes of computation while preserving the power of classical translocation mapping. Article summaryThe Drosophila melanogaster Y chromosome is difficult to study because it consists largely of repetitive, non-recombining DNA. Researchers have traditionally mapped Y-linked genes using slow, labor-intensive genetic crosses. Here we introduce Digital Kennison, a computational pipeline that replicates this classical mapping strategy using DNA sequence data instead of live flies. By comparing a query sequence against genomic databases built from fly strains, the pipeline assigns it to one of six Y-chromosome regions and reports a confidence score. Tested on 60 known sequences, it achieved 97% precision, corrected an annotation error, and mapped previously unplaced sequences in minutes rather than weeks.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Transposable element profiles reveal cell line identity and loss of heterozygosity in Drosophila cell culture 96%
- Improved Genotype Inference Reveals Cis- and Trans-Driven Variation in the Loss-of-Heterozygosity Rates in Yeast 94%
- Replication timing analysis in polyploid cells reveals Rif1 uses multiple mechanisms to promote underreplication in Drosophila. 94%
Similar papers in this journal
Similar papers in this journal
- A new high-quality genome assembly and annotation for the threatened Florida Scrub-Jay (Aphelocoma coerulescens) 95%
- Minimizing detection bias of somatic mutations in a highly heterozygous oak genome 95%
- A High-quality Oxford Nanopore Assembly of the Hourglass Dolphin (Lagenorhynchus cruciger) Genome 94%
Similar papers in this journal
Similar papers in this journal
- Chromonomer: a tool set for repairing and enhancing assembled genomes through integration of genetic maps and conserved synteny 95%
- A dense linkage map for a large repetitive genome: discovery of the sex-determining region in hybridising fire-bellied toads (Bombina bombina and B. variegata) 95%
- A telomere to telomere assembly of Oscheius tipulae and the evolution of rhabditid nematode chromosomes 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.