DomBpred: protein domain boundary predictor using inter-residue distance and domain-residue level clustering
Yu, Z.; Peng, C.; Liu, J.; Zhang, B.; Zhou, X.; Zhang, G.
Show abstract
Domain boundary prediction is one of the most important problems in the study of protein structure and function, especially for large proteins. At present, most domain boundary prediction methods have low accuracy and limitations in dealing with multi-domain proteins. In this study, we develop a sequence-based protein domain boundary predictor, named DomBpred. In DomBpred, the input sequence is firstly classified as either a single-domain protein or a multi-domain protein through a designed effective sequence metric based on a constructed single-domain sequence library. For the multi-domain protein, a domain-residue level clustering algorithm inspired by Ising model is proposed to cluster the spatially close residues according inter-residue distance. The unclassified residues and the residues at the edge of the cluster are then tuned by the secondary structure to form potential cut points. Finally, a domain boundary scoring function is proposed to recursively evaluate the potential cut points to generate the domain boundary. DomBpred is tested on a large-scale test set of FUpred comprising 2549 proteins. Experimental results show that DomBpred better performs than the state-of-the-art methods in classifying whether protein sequences are composed by single or multiple domains, and the Matthews correlation coefficient is 0.882. Moreover, on 849 multi-domain proteins, the domain boundary distance and normalised domain overlap scores of DomBpred are 0.523 and 0.824, respectively, which are 5.0% and 4.2% higher than those of the best comparison method, respectively. Comparison with other methods on the given test set shows that DomBpred outperforms most state-of-the-art sequence-based methods and even achieves better results than the top-level template-based method.
Matching journals
The top 4 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- Rprot-Vec: A deep learning approach for fast protein structure similarity calculation 98%
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 97%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 97%
Similar papers in this journal
- Sequence alignment using machine learning for accurate template-based protein structure prediction 97%
- DELPHI: accurate deep ensemble model for protein interaction sites prediction 97%
- A de novo protein structure prediction by iterative partition sampling, topology adjustment, and residue-level distance deviation optimization 97%
Similar papers in this journal
- Elucidation of Genome-wide Understudied Proteins targeted by PROTAC-induced degradation using Interpretable Machine Learning 96%
- Pathfinder: protein folding pathway prediction based on conformational sampling 96%
- GCNCDA: A New Method for Predicting CircRNA-Disease Associations Based on Graph Convolutional Network Algorithm 95%
Similar papers in this journal
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
- Navigating the Unstructured by Evaluating AlphaFold's Efficacy in Predicting Missing Residues and Structural Disorder in Proteins 95%
- Using deep maxout neural networks to improve the accuracy of function prediction from protein interaction networks 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.