DeepPrime2Sec: Deep Learning for Protein Secondary Structure Prediction from the Primary Sequences
Asgari, E.; Poerner, N.; McHardy, A.; Mofrad, M.
Show abstract
MotivationHere we investigate deep learning-based prediction of protein secondary structure from the protein primary sequence. We study the function of different features in this task, including one-hot vectors, biophysical features, protein sequence embedding (ProtVec), deep contextualized embedding (known as ELMo), and the Position Specific Scoring Matrix (PSSM). In addition to the role of features, we evaluate various deep learning architectures including the following models/mechanisms and certain combinations: Bidirectional Long Short-Term Memory (BiLSTM), convolutional neural network (CNN), highway connections, attention mechanism, recurrent neural random fields, and gated multi-scale CNN. Our results suggest that PSSM concatenated to one-hot vectors are the most important features for the task of secondary structure prediction.\n\nResultsUtilizing the CNN-BiLSTM network, we achieved an accuracy of 69.9% and 70.4% using ensemble top-k models, for 8-class of protein secondary structure on the CB513 dataset, the most challenging dataset for protein secondary structure prediction. Through error analysis on the best performing model, we showed that the misclassification is significantly more common at positions that undergo secondary structure transitions, which is most likely due to the inaccurate assignments of the secondary structure at the boundary regions. Notably, when ignoring amino acids at secondary structure transitions in the evaluation, the accuracy increases to 90.3%. Furthermore, the best performing model mostly mistook similar structures for one another, indicating that the deep learning model inferred high-level information on the secondary structure.\n\nAvailabilityThe developed software called DeepPrime2Sec and the used datasets are available at http://llp.berkeley.edu/DeepPrime2Sec.\n\nContactmofrad@berkeley.edu
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Improving sequence-based modeling of protein families using secondary structure quality assessment 97%
- Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function 97%
- A Gated Graph Transformer for Protein ComplexStructure Quality Assessment and its Performancein CASP15 96%
Similar papers in this journal
- SPDesign: protein sequence designer based on structural sequence profile using ultrafast shape recognition 96%
- A Unified Protein Embedding Model with Local and Global Structural Sensitivity 96%
- EGRET: Edge Aggregated Graph Attention Networks and Transfer Learning Improve Protein-Protein Interaction Site Prediction 95%
Similar papers in this journal
- Flattening the curve - How to get better results with small deep-mutational-scanning datasets 96%
- DisCovER: distance- and orientation-based covariational threading for weakly homologous proteins 95%
- Improving protein tertiary structure prediction by deep learning and distance prediction in CASP14 95%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 97%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- Estimating Protein Complex Model Accuracy Using Graph Transformers and Pairwise Similarity Graphs 94%
Similar papers in this journal
- RNA secondary structure prediction with Convolutional Neural Networks 95%
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.