Back

Whole-genome screening for near-diagnostic genetic markers for white oak species identification in Europe

KREMER, A.; Delcamp, A.; Lesur, I.; Wagner, S.; Christian, R.; Guichoux, E.; Leroy, T.

2023-12-01 plant biology
10.1101/2023.11.29.568959 bioRxiv
Show abstract

ContextIdentifying species in the European white oak complex has been a long standing concern in taxonomy, evolution, forest research and management. Quercus petraea, Q. robur, Q. pubescens and Q. pyrenaica are part of this species complex in western temperate Europe and hybridize in mixed stands, challenging species identification. AimsOur aim was to identify diagnostic single nucleotide polymorphisms (SNPs) for each of the four species that are suitable for routine use and rapid diagnosis in research and applied forestry. MethodsWe first scanned existing whole-genome and target-capture data sets in a reduced number of samples (training set) to identify candidate diagnostic SNPs, ie genomic positions being characterized by a reference allele in one species and by the alternative allele in all other species. Allele frequencies of the candidates SNPs were then explored in a larger, range-wide sample of populations in each species (validation step). ResultsWe found a subset of 38 SNPs (ten for Q. petraea, seven for Q. pubescens, nine for Q. pyrenaica and twelve for Q. robur) that showed near-diagnostic features across their species distribution ranges with Q. pyrenaica and Q. pubescens exhibiting the highest and lowest diagnosticity, respectively. ConclusionsWe provide a new, efficient and reliable molecular tool for the identification of the species Q. petraea, Q. robur, Q. pubescens and Q. pyrenaica, which can be used as a routine tool in forest research and management. This study highlights the resolution offered by whole-genome sequencing data to design diagnostic marker sets for taxonomic assignment, even for species complexes with relatively low differentiation.

Matching journals

The top 7 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.