ProtTox: Toxin identification from Protein Sequences
Datta, D.; Muthiah, S.; Butler, P.; Islam, M. R.; Warren, A.; Ramakrishnan, N.
Show abstract
Toxin classification of protein sequences is a challenging task with real world applications in healthcare and synthetic biology. Due to an ever expanding database of proteins and the inordinate cost of manual annotation, automated machine learning based approaches are crucial. Approaches need to overcome challenges of homology, multi-functionality, and structural diversity among proteins in this task. We propose a novel deep learning based method ProtTox, that aims to address some of the shortcomings of previous approaches in classifying proteins as toxins or not. Our method achieves a performance of 0.812 F1-score which is about 5% higher than the closest performing baseline.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 94%
- Predicting RNA Sequence-Structure Likelihood via Structure-Aware Deep Learning 94%
Similar papers in this journal
- PRIEST - Predicting viral mutations with immune escape capability of SARS-CoV-2 using temporal evolutionary information 96%
- CRISPR-DIPOFF: An Interpretable Deep LearningApproach for CRISPR Cas-9 Off-Target Prediction 95%
- Interpretable and Generalizable Attention-Based Model for Predicting Drug-Target Interaction Using 3D Structure of Protein Binding Sites: SARS-CoV-2 Case Study and in-Lab Validation 95%
Similar papers in this journal
- Building explainable graph neural network by sparse learning for the drug-protein binding prediction 94%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 93%
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.