Predicting Residues Involved in Anti-DNA Autoantibodies with Limited Neural Networks
St Clair, R.; Teti, M.; Pavlovic, M.; Hahn, W.; Barenholtz, E.
Show abstract
Computer-aided rational vaccine design (RVD) and synthetic pharmacology are rapidly developing fields that leverage existing datasets for developing compounds of interest. Computational proteomics utilizes algorithms and models to probe proteins for functional prediction. A potentially strong target for such a computational approach is autoimmune antibodies which are the result of broken tolerance in the immune system where it cannot distinguish "self" from "non-self" resulting in attack of its own structures (proteins and DNA, mainly). The information on structure, function and pathogenicity of autoantibodies may assist in engineering RVD against autoimmune diseases. Current computational approaches exploit large datasets curated with extensive domain knowledge, most of which include the need for many computational resources and have been applied indirectly to problems of interest for DNA, RNA, and monomer protein binding. Here, we present a novel method for discovering potential binding sites. We employed long short-term memory (LSTM) models trained on FASTA primary sequences directly to predict protein binding in DNA-binding hydrolytic antibodies (abzymes). We also employed CNN models applied to the same dataset. While the CNN model outperformed the LSTM on the primary task of binding prediction, analysis of internal model representations of both models showed that the LSTM models highlighted sub-sequences that were more strongly correlated with sites known to be involved in binding. These results demonstrate that analysis of internal processes of recurrent neural network models may serve as a powerful tool for primary sequence analysis.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 95%
- SpatialPPI: three-dimensional space protein-protein interaction prediction with AlphaFold Multimer 94%
- DrugForm-DTA: Towards real-world drug-target binding Affinity Model 94%
Similar papers in this journal
Similar papers in this journal
- Employing Machine Learning Techniques to Detect Protein-Protein Interaction: A Survey, Experimental, and Comparative Evaluations 95%
- AE-LGBM: Sequence-Based Novel Approach To Detect Interacting Protein Pairs via Ensemble of Autoencoder and LightGBM. 94%
- ToxinPred 3.0: An improved method for predicting the toxicity of peptides 94%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 94%
Similar papers in this journal
- DBpred: A deep learning method for the prediction of DNA interacting residues in protein sequences 95%
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 94%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.