EPEPDI: prediction of binding free energy changes from missense mutations in double and single-stranded DNA-binding proteins
YU, X.
Show abstract
Predicting changes in binding free energy due to missense mutations (MMs) in protein-DNA interactions (PDIs) is vital for understanding disease mechanisms and advancing therapeutic strategies. However, many existing models fail to account for the unique characteristics of MMs in double-stranded DNA binding proteins (DSBs) and single-stranded DNA binding proteins (SSBs). To address this, we constructed a comprehensive dataset from diverse sources, clearly distinguishing between DSBs and SSBs. Using sequence-based embeddings from pre-trained protein language models, including ESM2, ProtTrans, and ESM1v, we developed EPEPDI, a deep learning framework that integrates these embeddings through a multi-channel architecture. To refine predictive accuracy, we introduced an information entropy-based algorithm, determining 181 residues as the optimal sequence length where amino acid contributions and entropy dynamics balance. This approach boosts both precision and computational efficiency, enabling scalable analysis of mutation impacts on DNA-binding proteins. Ablation studies validated optimal feature combinations, demonstrating that EPEPDI outperforms existing approaches, achieving an average Pearson correlation coefficient of 0.755 on the MPD276 dataset via ten-fold cross-validation and 0.632 on independent tests for both DSBs and SSBs. This work highlights the importance of distinguishing DSBs and SSBs in PDIs and shows the potential of advanced machine learning in biological research.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- From Proteins to Ligands: Decoding Deep Learning Methods for Binding Affinity Prediction 96%
- ProAffinity-GNN: A Novel Approach to Structure-based Protein-Protein Binding Affinity Prediction via a Curated Dataset and Graph Neural Networks 96%
- From Signal to Symphony: Exploring 2D Sequence Representations for Protein Function Prediction 96%
Similar papers in this journal
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 96%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 95%
- Data-efficient protein mutational effect prediction with weak supervision by molecular simulation and protein language models 95%
Similar papers in this journal
Similar papers in this journal
- Enhancing predictions of protein stability changes induced by single mutations using MSA-based Language Models 97%
- Pair-EGRET: enhancing the prediction of protein-proteininteraction sites through graph attention networks and protein language models 95%
- Attention-based approach to predict drug-target interactions across seven target superfamilies 95%
Similar papers in this journal
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 95%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
- Improving classification of correct and incorrect protein-protein docking models by augmenting the training set 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.