De novo whole-genome assembly and annotation of a high-quality coffee variety from the primary origin of coffee, Coffea arabica var. Geisha
Medrano, J. F.; Cantu, D.; Minio, A.; Dreischer, C.; Gibbons, T.; Chin, J.; Chen, S.; Van Deynze, A.; Hulse-Kemp, A. M.
Show abstract
Geisha coffee is recognized for its unique aromas and flavors and accordingly, has achieved the highest prices in the specialty coffee markets. We report the development of a chromosome-level, well-annotated, genome assembly of Coffea arabica var. Geisha, considered an Ethiopian landrace thatrepresents germplasm from the Ethiopian center of origin of coffee. We used a hybrid de novo assembly approach combining two long-reads single molecule sequencing technologies, Oxford Nanopore and Pacific Biosciences, together with scaffolding with Hi-C libraries. The final assembly is 1.03GB in size with BUSCO assessment of the assembly completeness of 97.7% of single-copy orthologs clusters. RNAseq and IsoSeq data were used as transcriptional experimental evidence for annotation and gene prediction revealing the presence of 47,062 gene loci encompassing 53,273 protein-coding transcripts. Comparison of the assembly to the progenitor subgenomes, separated the set of chromosome sequences inherited from C. canephora from those of C. eugenioides., Corresponding orthologs between Geisha and Red Bourbon had a 99.67% median identity, higher than what we observe with the progenitor assemblies (median 97.28%). Both, Geisha and Red Bourbon contain an inversion on Chromosome 10 relative to the pseudomolecules of the genetic material inherited from the two progenitors that must have happened before the separation in the geographical migration of the two varieties. Lending support of a single allopolyploidization event that gave origin to C. arabica after the hybridization event with the two progenitor lines. Broadening the availability of high-quality genome assemblies of Coffea arabica varieties, paves the way for understanding the evolution and domestication of coffee, as well as the genetic basis and environmental interactions of why a variety like Geisha is capable of producing beans with such exceptional and unique high-quality.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- De novo whole-genome assembly of Chrysanthemum makinoi, a key wild ancestor to hexaploid Chrysanthemum 98%
- Genome assembly, annotation and comparative analysis of the cattail Typha latifolia 97%
- Diploid chromosome-scale assembly of the Muscadinia rotundifolia genome supports chromosome fusion and disease resistance gene expansion during Vitis and Muscadinia divergence 97%
Similar papers in this journal
- Chromosome-level genome assemblies for two quinoa inbred lines from northern and southern highlands of Altiplano where quinoa originated 98%
- A Partially Phase-Separated Genome Sequence Assembly of the Vitis Rootstock 'Börner' (Vitis riparia x Vitis cinerea) and its Exploitation for Marker Development and Targeted Mapping 96%
- Oat chromosome and genome evolution defined by widespread terminal intergenomic translocations in polyploids 95%
Similar papers in this journal
- The Genome of Chenopodium ficifolium: Developing Genetic Resources and a Diploid Model System for Allotetraploid Quinoa 98%
- Conserving a threatened North American walnut: a chromosome-scale reference genome for butternut (Juglans cinerea) 97%
- A haplotype-complete chromosome-level assembly of octoploid Urochloa humidicola cv. Tully reveals multiple genomic compositions and evolutionary histories in the species 97%
Similar papers in this journal
- A view of the pan-genome of domesticated cowpea (Vigna unguiculata Walp.) 97%
- New insights into homoeologous copy number variations in the hexaploid wheat genome 95%
- Chromosome-level haplotype-resolved genome assembly provides insights into the highly heterozygous genome of Italian ryegrass (Lolium multiflorum Lam.) 95%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.