Back

Evolutionary fingerprinting of protein-coding genes in RNA viruses

Munoz-Baena, L.; Castelan-Sanchez, H.; Bagherichimeh, S.; Magbor, P.; Rojas-Vargas, J.; Khan, A.; Olabode, A. S.; Poon, A.

2025-12-06 evolutionary biology
10.64898/2025.12.05.692645 bioRxiv
Show abstract

RNA viruses evolve rapidly to adapt to changing host environments. Much of this adaptation occurs at the level of protein-coding genes. Thus, some of the most well-characterized examples of rapid adaptation have been found in virus proteins that are exposed on the surface of the viral particle, where they mediate host receptor binding and cell entry. To investigate whether surface-exposed proteins and other proteins encoded by viruses exhibit different patterns of evolution under selection, we analyzed 244 protein-coding genes from 28 species of RNA viruses representing 15 taxonomic families. First, we show that gene-wide rates of non-synonymous (dN) and synonymous (dS) substitutions do not differentiate between categories of proteins. To provide a more detailed comparison between genes, we inferred for each alignment the bivariate posterior distribution over a fixed grid of codon site-specific dN and dS values. This distribution is the genes evolutionary fingerprint. Next, we computed the Wasserstein distance for every pair of fingerprints, which is analogous to amount of work required to reshape one distribution to another. After compensating for differences in genetic variation among viruses and proteins, we found that surface-exposed proteins could not be distinguished from non-exposed proteins in the space induced by the Wasserstein distance matrix. However, surface-exposed proteins from enveloped viruses were significantly clustered apart from their counterparts in non-enveloped viruses. In contrast, there was no significant separation between these categories of viruses for proteins with polymerase activity. We show that this pattern is more consistent with relaxed purifying selection than adaptive evolution in proteins associated with viral envelopes.

Matching journals

The top 1 journal accounts for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.