Back

A chromosome-scale genome assembly of the Swiss Lolium multiflorum ecotype Tremona reveals a scalable method to purge spurious duplications

Piat, L.; Herren, G.; Grieder, C.; Roulin, A. C.

2026-08-20 genomics
10.64898/2026.08.18.745395 bioRxiv
Show abstract

Italian ryegrass (Lolium multiflorum) is a key temperate forage species underpinning livestock production in Europe. Genomic resources remain limited by its large (2.2 Gb), repetitive, and highly heterozygous genome. Here, we present a high-quality chromosome-scale genome assembly of the Swiss L. multiflorum ecotype Tremona, collected in 2008 in Ticino, Switzerland, and subsequently incorporated into recurrent breeding cycles in the Swiss breeding program. To address systematic assembly artefacts caused by unresolved haplotypes in our initial PacBio HiFi assembly, we developed ParaLies, a post-assembly tool that identifies and removes artefactual duplications based on sequence divergence while preserving true paralogous gene copies. ParaLies reduced the duplicated BUSCO rate from 16.91% to 6.72% without loss of bona fide genomic content. The resulting assembly has a contig N50 of 15.69 Mb and captures 94% of the expected 2.2-Gb genome size. We further analyzed whole-genome resequencing data from Tremona, additional Swiss ecotypes, and publicly available North American germplasm. Tremona was genetically homogeneous, with no evidence of pronounced recent bottlenecks or substantial within-population structure, and was genetically distinct from the other Swiss ecotypes analyzed. Together, the Tremona genome and ParaLies provide valuable resources for L. multiflorum genomics and breeding and demonstrate a scalable approach for reducing haplotype-induced redundancy in highly heterozygous genomes.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.