Decoding Protein Aggregation through Computational Approach: Identification and Scoring of Aggregation-Prone Regions in Protein Sequences
Kaushik, R.; Launey, T.
Show abstract
Protein aggregation is a critical phenomenon associated with numerous neurodegenerative and systemic diseases. Understanding the propensity of proteins to aggregate is essential for unraveling the molecular basis of these disorders and for design and engineering of novel proteins or modulating the activity/stability of enzymatic proteins. Here, we present APR-Score, a novel machine-learning based computational method designed to identify aggregation-prone regions within protein sequences. ARP-Score leverages a combination of sequence-based features to predict regions of proteins that are prone to aggregate. The APR-Score harnessed the information ingrained in the compiled sequence and structural features to provide state-of-the-art accuracy. The APR-Score is assessed by conducting rigorous cross-validation experiments on the training dataset and further validated on an independent test dataset. The APR-Score prediction models demonstrated robustness and reliability in discriminating aggregation-prone regions from non-aggregating ones on an independent dataset, achieving Mathews correlation coefficient (MCC) 0.81, precision 0.89, and F1-Score 0.91. The APR-Score offers a valuable tool for researchers investigating protein aggregation-related diseases, as it can expedite the identification of aggregation-prone regions, aiding in the development of targeted therapies and diagnostic tools. The computational protein design and engineering regimes can be facilitated through APR-Score based identification and screening of aggregation prone protein sequences.
Matching journals
The top 7 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
- VISUALIZING GAUSSIAN-CHAIN LIKE STRUCTURAL MODELS OF HUMAN alpha-SYNUCLEIN IN MONOMERIC PRE-FIBRILLAR STATE: SOLUTION SAXS DATA AND MODELING ANALYSIS 95%
- PredIDR: Accurate prediction of protein intrinsic disorder regions using deep convolutional neural network 94%
- The Distal-Proximal Relationships Among the Human Moonlighting Proteins: Evolutionary hotspots and Darwinian checkpoints 94%
Similar papers in this journal
- HAIRpred: Prediction of human antibody interacting residues in an antigen from its primary structure 94%
- Convergent behavior of extended stalk regions from staphylococcal surface proteins with widely divergent sequence patterns 93%
- Target-template relationships in protein structure prediction and their effect on the accuracy of thermostability calculations 92%
Similar papers in this journal
- Peptides derived from gp43, the most antigenic protein from Paracoccidioides brasiliensis, form amyloid fibrils in vitro: implications for vaccine development 95%
- PACT - Prediction of Amyloid Cross-interaction by Threading 95%
- Predicting interchain contacts for homodimeric and homomultimeric protein complexes using multiple sequence alignments of monomers and deep learning 94%
Similar papers in this journal
- Sequence alignment using machine learning for accurate template-based protein structure prediction 94%
- A de novo protein structure prediction by iterative partition sampling, topology adjustment, and residue-level distance deviation optimization 94%
- A3D Database: Structure-based Protein Aggregation Predictions for the Human Proteome 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.