Back

Fully Phased Telomere-to-Telomere Assemblies for Thoroughbred Horse and Donkey Haplotypes derived from a Mule Illuminate the Peculiar Evolution of Equid Centromeres

Li, K.; Cappelletti, E.; Dessaix, C.; Ciosek, J. L.; Robyn, E. D.; Johnson, L. C.; Hussien, N. A.; Arias, X. R.; Adelson, D. L.; Raudsepp, T.; Piras, F. M.; Smith, M.; Hudson, E.; Pickett, B.; Koren, S.; Walenz, B.; Sison, C.; Crawford, J.; Bouffard, G.; Brooks, S.; Phillippy, A.; Miller, D.; Antczak, D. F.; Cullen, J.; Stroupe, S.; Davis, B.; McCue, M.; Durward-Akhurst, S.; Petersen, J. L.; Giulotto, E.; Kalbfleisch, T.

2026-02-27 molecular biology
10.64898/2026.02.26.707634 bioRxiv
Show abstract

We present telomere-to-telomere genome assemblies of a Thoroughbred horse and a donkey derived from their mule offspring. Now adopted and annotated by NCBI as reference genomes, these assemblies resolve previously inaccessible regions, including satellite arrays, duplications, and telomeres. Equids are known to exhibit an uncoupling between satellite DNA and centromeric function. The completeness of these assemblies enabled annotation of both satellite-based and satellite-free centromeres, as well as non-centromeric satellite loci, revealing notable centromeric plasticity. They also allowed detailed characterization of the variable binding domains of CENP-A--the epigenetic determinant of centromere identity--and CENP-B, whose association with CENP-A, previously considered typical based on a few model organisms, is absent in equids. Comparative analyses of satellite repeats and centromere positions provide new insights into the accelerated karyotypic reshuffling in equid evolution. These assemblies represent foundational resources for equid genomics and support ongoing initiatives such as the Equine Pangenome Project.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.