Nanopore-based genome assembly and the evolutionary genomics of basmati rice
Choi, J. Y.; Lye, Z.; Groen, S.; Dai, X.; Rughani, P.; Zaaijer, S.; Harrington, E.; Juul, S.; Purugganan, M.
10.1101/396515 bioRxivShow abstract
BACKGROUNDThe circum-basmati group of cultivated Asian rice (Oryza sativa) contains many iconic varieties and is widespread in the Indian subcontinent. Despite its economic and cultural importance, a high-quality reference genome is currently lacking, and the groups evolutionary history is not fully resolved. To address these gaps, we used long-read nanopore sequencing and assembled the genomes of two circum-basmati rice varieties, Basmati 334 and Dom Sufid.\n\nRESULTSWe generated two high-quality, chromosome-level reference genomes that represented the 12 chromosomes of Oryza. The assemblies showed a contig N50 of 6.32Mb and 10.53Mb for Basmati 334 and Dom Sufid, respectively. Using our highly contiguous assemblies we characterized structural variations segregating across circum-basmati genomes. We discovered repeat expansions not observed in japonica--the rice group most closely related to circum- basmati--as well as presence/absence variants of over 20Mb, one of which was a circum- basmati-specific deletion of a gene regulating awn length. We further detected strong evidence of admixture between the circum-basmati and circum-aus groups. This gene flow had its greatest effect on chromosome 10, causing both structural variation and single nucleotide polymorphism to deviate from genome-wide history. Lastly, population genomic analysis of 78 circum-basmati varieties showed three major geographically structured genetic groups: (1) Bhutan/Nepal group, (2) India/Bangladesh/Myanmar group, and (3) Iran/Pakistan group.\n\nCONCLUSIONAvailability of high-quality reference genomes from nanopore sequencing allowed functional and evolutionary genomic analyses, providing genome-wide evidence for gene flow between circum-aus and circum-basmati, the nature of circum-basmati structural variation, and the presence/absence of genes in this important and iconic rice variety group.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- A pangenome analysis pipeline (PSVCP) provides insights into rice functional gene identification 97%
- Automated assembly scaffolding elevates a new tomato system for high-throughput genome editing 95%
- Higher order repeat structures reflect diverging evolutionary paths in maize centromeres and knobs 95%
Similar papers in this journal
- Genome assembly and characterization of a complex zfBED-NLR gene-containing disease resistance locus in Carolina Gold Select rice with Nanopore sequencing 97%
- Genetic and environmental influences on the distributions of three chromosomal drive haplotypes in maize 94%
- Gene disruption by structural mutations drives selection in US rice breeding over the last century 94%
Similar papers in this journal
- A high-quality pseudo-phased genome for Melaleuca quinquenervia shows allelic diversity of NLR-type resistance genes 96%
- Long-reads assembly of the Brassica napus reference genome, Darmor-bzh 95%
- The draft nuclear genome assembly of Eucalyptus pauciflora: new approaches to comparing de novo assemblies 94%
Similar papers in this journal
- Evolutionary genomics of structural variation in Asian rice (Oryza sativa) and its wild progenitor (O. rufipogon) 93%
- Molecular Parallelism Underlies Convergent Highland Adaptation of Maize Landraces 93%
- Global patterns of subgenome evolution in organelle-targeted genes of sixallotetraploid angiosperms 93%
Similar papers in this journal
- Benchmarking long-read variant calling in diploid and polyploid genomes: insights from human and plants 95%
- A chromosome-level genome assembly of amodel conifer plant, the Japanese cedar,Cryptomeria japonica D. Don 95%
- Genomic architecture of 5S rDNA cluster and its variations within and between species 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.