Back

Universal single-copy genes and 16S rDNA present incongruent evolutionary histories in Vibrio

Garcia-Florez, A.; Leunda-Esnaola, A.; Arrufat, P.; Kaberdin, V. R.; Pearman, P. B.

2025-12-07 microbiology
10.64898/2025.12.07.692777 bioRxiv
Show abstract

A common technique for the study of the diversity and evolution of microbial communities is 16S rDNA sequencing. However, high sequence identity and variable copy number constrain the application of 16S rDNA in differentiation of closely related taxa and estimation of species relative abundance in environmental samples. A promising alternative is the use of universal single-copy genes (USCGs) as phylogenetic markers. We develop this by analyzing a set of USCG loci from the genus Vibrio, which holds over 100 species of substantial ecological and epidemiological relevance. The phylogenetic histories of these loci, of representative copies of 16S and 23S rDNA genes, and of a collection of 16S rDNA partial sequences were reconstructed using Bayesian inference. Taxon resolution was assessed according to consensus tree topology and clade credibility values. In addition, the congruence among posterior distributions of phylogenetic estimates of the different loci was calculated using Robinson-Foulds distances and visualized with non-metric multidimensional scaling (NMDS). Phylogenetic analyses reveal that USCG loci produce highly resolved trees in comparison to those of 16S and 23S rDNA sequences. We also observe relatively high congruence among phylogenies of USCG loci while rDNA phylogenies diverge from these. The loci mfd and uvrC are highlighted for further research on Vibrio evolution and analysis of environmental samples. Moreover, possible sources of phylogenetic incongruence between USCG and rDNA loci include differential susceptibility to horizontal gene transfer, as potentially explained by the complexity hypothesis, or lack of phylogenetic information due to limited sequence variability in rDNA sequences.

Matching journals

The top 6 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.