Back

Pervasive cryptic selection in the human noncoding genome

Ramesh, S.; Di, C.; Lohmueller, K.

2026-06-10 evolutionary biology
10.64898/2026.06.09.731256 bioRxiv
Show abstract

The prevailing dogma in evolutionary genetics holds that mutations within sequences that are conserved across a phylogeny are deleterious in those species, and mutations outside are neutrally evolving. Indeed, such comparative genomic approaches have estimated that mutations in approximately 5% of the human genome experience negative selection. However, sites that have biological function in certain lineages but not in others, i.e. functional turnover, may violate this assumption since these sites may be invisible to comparative genomic approaches. Thus, the extent of such cryptic, or hidden, negative selection remains elusive. Here, we developed a statistical test to detect cryptic selection in human polymorphism data. Applying our approach to simulated data shows that cryptic selection shapes the site frequency spectrum (SFS) and the statistical detection power depends on the proportion of mutations experiencing cryptic selection, the amount of sequence tested, and the sample size. We applied our method to polymorphism data from the 1000 Genomes Project, comparing variants in putatively functional noncoding regions to those in putatively neutral regions. We detected pervasive signals of cryptic selection in putatively functional regions, even after filtering out the top 70% of conserved sites. Using simulations with varying levels of cryptic selection, we estimated the extent of genome-wide constraint in the human genome. Our approximation suggests that mutations in at least 7% of the human genome are under negative selection, which is greater than the estimates from conservation-based methods, and that many of these mutations have escaped detection by comparative genomic methods. In sum, our results highlight the evolutionary dynamic nature of the noncoding genome and suggest the need to account for functional turnover when identifying putatively neutral variants for evolutionary analyses.

Matching journals

The top 2 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.