Predicting mutation-rate variation across the genome using epigenetic data
Katori, M.; Kobayashi, T. J.; Nordborg, M.; Shi, S.
Show abstract
Mutation rate variation is a fundamental driver of evolution, yet how it is locally patterned across genomes and structured by chromatin context remains unresolved. Here, we integrate genome-wide profiles of histone marks, DNA methylation and chromatin accessibility in Arabidopsis thaliana with de novo mutation data to model mutation probability at the level of coding sequence (CDS). Using non-negative matrix factorization, we identify 15 combinatorial epigenetic patterns whose graded mixtures stratify CDSs into six classes with distinct mutation probabilities. A generalized linear model based on pattern weights predicts local mutation probability and outperforms models based on sequence context, expression and classical genomic categories. These patterns capture context-dependent variation that is obscured by gene-level summaries and single-feature analyses. Cluster-level differences are partly retained in mutation-accumulation lines, indicating persistence into heritable mutational input. Under hypoxia, stress-responsive chromatin remodeling redistributes epigenetic contexts associated with higher predicted mutation probability toward hypoxia-responsive genes and DNA-repair pathways. Together, our results provide a CDS-resolved and interpretable framework linking combinatorial epigenomic context to mutational input, clarifying how dynamic chromatin states shape local mutation-rate heterogeneity.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Chromatin Regulates Bipartite-Classified Small RNA Expression to Maintain Epigenome Homeostasis in Arabidopsis 95%
- INFIMA leverages multi-omics model organism data to identify effector genes of human GWAS variants 95%
- Co-opted transposons help perpetuate conserved higher-order chromosomal structures 94%
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.