AAclust: k-optimized clustering for selecting redundancy-reduced sets of amino acid scales
Breimann, S.; Frishman, D.
Show abstract
SummaryAmino acid scales are crucial for sequence-based protein prediction tasks, yet no gold standard scale set or simple scale selection methods exist. We developed AAclust, a wrapper for clustering models that require a pre-defined number of clusters k, such as k-means. AAclust obtains redundancy-reduced scale sets by clustering and selecting one representative scale per cluster, where k can either be optimized by AAclust or defined by the user. The utility of AAclust scale selections was assessed by applying machine learning models to 24 protein benchmark datasets. We found that top-performing scale sets were different for each benchmark dataset and significantly outperformed scale sets used in previous studies. Notably, model performance showed a strong positive correlation with the scale set size. AAclust enables a systematic optimization of scale-based feature engineering in machine learning applications. Availability and implementationThe AAclust algorithm is part of AAanalysis, a Python-based framework for interpretable sequence-based protein prediction, which will be made freely accessible in a forthcoming publication. ContactStephan Breimann (Stephan.Breimann@dzne.de) and Dmitrij Frishman (dimitri.frischmann@tum.de) Supplementary informationFurther details on methods and results are provided in Supplementary Material.
Matching journals
The top 3 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- SPRI: Structure-Based Pathogenicity Relationship Identifier for Predicting Effects of Single Missense Variants and Discovery of Higher-Order Cancer Susceptibility Clusters of Mutations 94%
- PLMFit : Benchmarking Transfer Learning with Protein Language Models for Protein Engineering 94%
- An Analysis of Protein Language Model Embeddings for Fold Prediction 94%
Similar papers in this journal
- Zero-shot segmentation using embeddings from a protein language model identifies functional regions in the human proteome 95%
- From complete cross-docking to partners identification and binding sites predictions 94%
- Hybridized distance- and contact-based hierarchical structure modeling for folding soluble and membrane proteins 94%
Similar papers in this journal
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 94%
- TemBERTure: Advancing protein thermostability prediction with Deep Learning and attention mechanisms 93%
- DeepRank-GNN-esm: A Graph Neural Network for Scoring Protein-Protein Models using Protein Language Model 93%
Similar papers in this journal
- Towards mechanistic models of mutational effects: Deep Learning on Alzheimer's Aβ peptide 93%
- Systematic Investigation of Machine Learning on Limited Data: A Study on Predicting Protein-Protein Binding Strength 92%
- G-PLIP: Knowledge graph neural network for structure-free protein-ligand bioactivity prediction 92%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.