DPAC: Prediction and Design of Protein-DNA Interactions via Sequence-Based Contrastive Learning
Chen, L. T.; Pulugurta, R.; Vure, P.; Chatterjee, P.
Show abstract
Interactions between DNA and proteins are pivotal in natural biological processes, and designing proteins that can bind to DNA with high specificity is crucial for advancing genomic technologies. Existing state-of-the-art models for both modeling and designing protein-DNA interactions primarily rely on structural information, facing limitations in scalability and efficiency for large-scale applications. Notable methods like AlphaFold 3 and RosettaTTAFold All-Atom exist, but they are inefficient and inherently struggle at modeling conformationally unstable proteins, such as transcription factors, which arguably represent the most important class of DNA-binding proteins. Here, we present DPAC1 (DNA-Protein binding Alignment via Contrastive learning), which leverages pre-trained protein and DNA language models via a contrastive loss to align the two modalities in a high-dimensional shared latent space. DPAC not only significantly accelerates the design process compared to current structure-based methods but also demonstrates a strong ability to differentiate real binders from non-binders. Our model achieves an AUC score of 0.591 on a low identity set, outperforming state-of-the-art structure-based methods. Additionally, DPAC integrates simulated annealing for the design of new protein sequences with optimized DNA binding affinity, successfully recovering binding affinity in engineered sequences by up to 20% in in silico tests. Our results highlight DPACs potential for facilitating the design and discovery of sequence-specific DNA-binding proteins, paving the way for advancements in genomic research and biotechnology applications.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DNAffinity: A MACHINE-LEARNING APPROACH TO PREDICT DNA BINDING AFFINITIES OF TRANSCRIPTION FACTORS 96%
- Predicting the DNA binding specificity of transcription factor mutants usingfamily-level biophysically interpretable machine learning 95%
- Mutation rate variations in the human genome are encoded in DNA shape 95%
Similar papers in this journal
- Multiple protein-DNA interfaces unravelled by evolutionary information, physico-chemical and geometrical properties 95%
- Predicting changes in protein thermodynamic stability upon point mutation with deep 3D convolutional neural networks 95%
- An Associative Memory Hamiltonian Model for DNA and Nucleosomes 95%
Similar papers in this journal
- Enzyme structure correlates with variant effect predictability 95%
- Multiscale Molecular Modelling of Chromatin with MultiMM: From Nucleosomes to the Whole Genome 93%
- A Multi-Layered Computational Structural Genomics Approach Enhances Domain-Specific Interpretation of Kleefstra Syndrome Variants in EHMT1 92%
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.