Network-based Machine Learning Approach for Structural Domain Identification in Proteins
Tiwari, A.; Parekh, N.
Show abstract
In the era of structural genomics, with a large number of protein structures becoming available, identification of domains is an important problem in protein function analysis as it forms the first step in protein classification. Domain identification has been an active area of research for over four decades and a wide range of automated methods have been proposed. In the proposed network-based machine learning approach, NML-DIP, a combination of supervised (SVM) and unsupervised (k-means) machine learning techniques are used for domain identification in proteins. The algorithm proceeds by first representing protein structure as a protein contact network and using topological properties, viz., length, density, and interaction strength (that assesses inter- and intra-domain interactions) as feature vectors in the first SVM to distinguish between single and multi-domain proteins. A second SVM is used to identify the number of domains in multi-domain proteins. Thus, it does not require a prior information of the number of domains. The k-means algorithm is then used to identify the domain boundaries that are assessed using CATH annotation. The performance of the proposed algorithm is evaluated on four benchmark datasets and compared with four state-of-the-art domain identification methods. The performance of the approach is comparable to other domain identification tools and works well even when the domains are formed with non-contiguous segments. The performance of the program is significantly improved for prior information about the number of domains. The algorithm is available at: https://bit.ly/NML-DIP.
Matching journals
The top 9 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- RAFTS3G - An efficient and versatile clustering software to analyses in large protein datasets 96%
- Rprot-Vec: A deep learning approach for fast protein structure similarity calculation 96%
- Binding affinity prediction for protein-ligand complex using deep attention mechanism based on intermolecular interactions 95%
Similar papers in this journal
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 96%
- Classification of protein binding ligands using structural dispersion of binding site atoms from principal axes 95%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
Similar papers in this journal
Similar papers in this journal
- Predicting interchain contacts for homodimeric and homomultimeric protein complexes using multiple sequence alignments of monomers and deep learning 95%
- A Convolution Based Computational Approach Towards DNA N6-methyladenine Site Identification and Motif Extraction in Rice Genome 94%
- Designing of thermostable proteins with a desired melting temperature 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.