Back

Population-level super-pangenome reveals genome evolution and empowers precision breeding in watermelon

Sun, H.; Zhang, J.; Liao, S.; Guo, S.; Zhou, Z.; Zhao, X.; Wu, S.; Zhao, J.; Gong, G.; Wang, J.; Li, M.; Yu, Y.; Ren, Y.; Tian, S.; Li, S.; Zhang, H.; Hammar, S. A.; McGregor, C.; Jarret, R.; Wechter, P.; Branham, S. E.; Kousik, C.; Levi, A.; Grumet, R.; Xu, Y.; Fei, Z.

2025-07-27 plant biology
10.1101/2025.07.25.666869 bioRxiv
Show abstract

Pangenomes are increasingly critical for harnessing crop genetic diversity, yet their resolution and utility are often limited by insufficient sampling of high-quality genome assemblies. Here, we report a population-level watermelon super-pangenome constructed from 138 reference-grade assemblies, including 135 newly generated near-gapless genomes representing all seven watermelon species. The super-pangenome captures approximately one million structural variants (SVs), enabling accurate variant genotyping across ~900 watermelon accessions and substantially expanding variant discovery both across and within species. Broader sampling within the pangenome provides insights into genome evolution among watermelon species and sheds light on the origin of cultivated watermelon. SV-inclusive genome-wide association studies enhance trait mapping resolution and identify a copy number variation upstream of ClFCI1 that regulates flesh color intensity in a dosage-dependent manner. Leveraging this comprehensive variation map, we developed high-accuracy genomic prediction models for 18 agronomic traits. Together, our findings and genomic resources establish a foundational framework for dissecting complex traits and accelerating precision breeding in watermelon, while offering a valuable model for SV-resolved pangenomics in crop species.

Published in Nature Genetics (predicted rank #19) · training set

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.