Back

Purifying selection acts on germline methylation to modify the CpG mutation rate at promoters

Boukas, L.; Bjornsson, H. T.; Hansen, K. D.

2020-07-04 evolutionary biology
10.1101/2020.07.04.187880 bioRxiv
Show abstract

In the study of molecular features like epigenetic marks, it is appealing to ascribe observed patterns to the action of natural selection. However, this conclusion requires a test for selection based on a well-defined notion of neutrality. Here, focusing on epigenetic marks at gene loci, we formalize what it means for an epigenetic mark to be neutral, and develop a test for selection. Our test respects the foundational aspect of epigenetics: trans-regulation by transcription factors and chromatin modifiers. It also enables adjustment for confounders. We establish that promoter DNA methylation, promoter H3K4me3 and exonic H3K36me3 are all under selection, and that this is unlikely to be a passive consequence of selection on gene expression. The effect of these epigenetic marks on fitness is arguably partly explained by a causal involvement in gene regulation. However, we also investigate the complementary explanation that DNA methylation and H3K36me3 are under selection in part because they modify the mutation rate of important genomic regions. We show that this explanation is consistent with empirical observations as well as population genetics theory, because of the trans-regulation. Exemplifying the protection of important regions from high mutability, we demonstrate that in humans the more intolerant to loss-of-function mutations a gene is, the lower its coding mutation rate is. Our framework for selection inference is simple but general, and we speculate that its core ideas will be useful for additional molecular features beyond epigenetic marks.

Matching journals

The top 3 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.