Back

Predicting mutation-rate variation across the genome using epigenetic data

Katori, M.; Kobayashi, T. J.; Nordborg, M.; Shi, S.

2026-02-03 bioinformatics
10.64898/2026.01.30.702885 bioRxiv
Show abstract

Mutation rate variation is a fundamental driver of evolution, yet how it is locally patterned across genomes and structured by chromatin context remains unresolved. Here, we integrate genome-wide profiles of histone marks, DNA methylation and chromatin accessibility in Arabidopsis thaliana with de novo mutation data to model mutation probability at the level of coding sequence (CDS). Using non-negative matrix factorization, we identify 15 combinatorial epigenetic patterns whose graded mixtures stratify CDSs into six classes with distinct mutation probabilities. A generalized linear model based on pattern weights predicts local mutation probability and outperforms models based on sequence context, expression and classical genomic categories. These patterns capture context-dependent variation that is obscured by gene-level summaries and single-feature analyses. Cluster-level differences are partly retained in mutation-accumulation lines, indicating persistence into heritable mutational input. Under hypoxia, stress-responsive chromatin remodeling redistributes epigenetic contexts associated with higher predicted mutation probability toward hypoxia-responsive genes and DNA-repair pathways. Together, our results provide a CDS-resolved and interpretable framework linking combinatorial epigenomic context to mutational input, clarifying how dynamic chromatin states shape local mutation-rate heterogeneity.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.