Back

mIDEA: An Interpretable Structure-Sequence Model for Methylation-Dependent Protein-DNA Binding Sensitivity

Zhang, Y.; Li, R.; Zhong, J.; Lin, X.

2025-11-16 bioinformatics
10.1101/2025.11.14.688575 bioRxiv
Show abstract

DNA methylation dynamically reshapes protein-DNA interaction landscapes, yet the mechanistic principles governing methylation-dependent recognition remain incompletely understood. Currently, there is no predictive framework that integrates methylation-aware binding specificity with 3D structure-based residue interactions for quantitatively assessing the impact of DNA methylation on protein-DNA interactions. To bridge this gap, we introduce mIDEA (Methylation-informed Interpretable protein-DNA Energy Associative model), a structure-aware, residue-level biophysical framework that predicts and interprets protein-methylated-DNA interactions by optimizing an amino acid-nucleotide energy matrix explicitly accounting for the modulatory effect of 5-methylcytosine. We validate mIDEA across diverse human transcription factors with distinct CpG methylation preferences. Using only structural templates and prior biochemical knowledge, mIDEA accurately recapitulates both "methyl-plus" and "methyl-minus" effects of DNA methylation on protein-DNA binding. Incorporating quantitative methylation-sensitive binding data further enables mIDEA to capture position-specific methylation effects. When integrated with whole-genome methylation profiles, the model enhances in vivo binding-site prediction by reducing false positives while preserving sensitivity compared to models trained on unmethylated sequences. Overall, mIDEA provides a principled framework for elucidating the molecular mechanisms underlying methylation-modulated protein-DNA recognition. The mIDEA source code is available at: https://github.com/LinResearchGroup-NCSU/IDEA_DNA_Methylation.

Matching journals

The top 5 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.