Back

An improved gray whale assembly highlights how allospecific reference-genome choice can affect genomic diversity estimates

Bruniche-Olsen, A.; Black, A.; Walden, K. K. O.; Fields, C. J.; Bickham, J. W.; DeWoody, J. A.

2025-01-24 evolutionary biology
10.1101/2025.01.23.634050 bioRxiv
Show abstract

Gray whales (Eschrichtius robustus) are unique as bottom feeding baleen whales and they have long been a conservation concern on both sides of the Pacific, in part because they migrate and disperse farther than any other species on earth. They experienced drastic population size declines due to environmental changes and commercial whaling. Here, we present an improved genome assembly for the gray whale. This genome assembly covers 2.4 Gb divided across 2689 contigs with an N50 of 15Mb. From the new assembly, we identify 75Mb sex-linked contigs and a identify 94.6% of searched genes based on Benchmarking Universal Single-Copy Ortholog score. We use the gray whale assembly to explore the effects of mapping to conspecific vs allospecific reference genomes when estimating genome-wide heterozygosity (H) and runs of homozygosity (ROH). The use of allospecific genomes significantly underestimate both H and ROH burden regardless of genomic distance and assembly quality. Our analyses highlight the importance of using contiguous conspecific assemblies in whale genomics and conservation. SignificanceGray whales are unique in their behavior, morphology and ecology. The novel and contiguous genome assembly presented here will serve as a valuable resource for studies of their population and comparative genomics, as well as for identifying key adaptations that have evolved in this clade. Finally, our results demonstrate that biases can arise when using allopatric assemblies to evaluate diversity metrics, so they should be used and interpreted with caution.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.