Back

The genome and population genomics of allopolyploid Coffea arabica reveal the diversification history of modern coffee cultivars

Salojarvi, J.; Rambani, A.; Yu, Z.; Guyot, R.; Strickler, S.; Lepelley, M.; Wang, C.; Rajaraman, S.; Rastas, P.; Zheng, C.; Munoz, D. S.; Meidanis, J.; Paschoal, A. R.; Bawin, Y.; Krabbenhoft, T.; Wang, Z. Q.; Fleck, S.; Aussel, R.; Bellanger, L.; Charpagne, A.; Fournier, C.; Kassam, M.; Lefebvre, G.; Metairon, S.; Moine, D.; Rigoreau, M.; Stolte, J.; Hamon, P.; Couturon, E.; Tranchant-Dubreuil, C.; Mukherjee, M.; Lan, T.; Engelhardt, J.; Stadler, P.; DeLemos, S. C.; Suzuki, S. I.; Sumirat, U.; ChingMan, W.; Dauchot, N.; Orozco-Arias, S.; Garavito, A.; Kiwuka, C.; Musoli, P.; Nalukenge, A.; Gu

2023-09-06 genomics
10.1101/2023.09.06.556570 bioRxiv
Show abstract

Coffea arabica, an allotetraploid hybrid of C. eugenioides and C. canephora, is the source of approximately 60% of coffee products worldwide, and its cultivated accessions have undergone several population bottlenecks. We present chromosome-level assemblies of a di-haploid C. arabica accession and modern representatives of its diploid progenitors, C. eugenioides and C. canephora. The three species exhibit largely conserved genome structures between diploid parents and descendant subgenomes, with no obvious global subgenome dominance. We find evidence for a founding polyploidy event 350,000-610,000 years ago, followed by several pre-domestication bottlenecks, resulting in narrow genetic variation. A split between wild accessions and cultivar progenitors occurred [~]30.5 kya, followed by a period of migration between the two populations. Analysis of modern varieties, including lines historically introgressed with C. canephora, highlights their breeding histories and loci that may contribute to pathogen resistance, laying the groundwork for future genomics-based breeding of C. arabica.

Matching journals

The top 4 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.