ProtParts, an automated web server for clustering and partitioning protein datasets
Li, Y.; Barra, C.
Show abstract
Data leakage originating from protein sequence similarity shared among train and test sets can result in model overfitting and overestimation of model performance and utility. However, leakage is often subtle and might be difficult to eliminate. Available clustering tools often do not provide completely independent partitions, and in addition it is difficult to assess the statistical significance of those differences. In this study, we developed a clustering and partitioning tool, ProtParts, utilizing the E-value of BLAST to compute pairwise similarities between each pair of proteins and using a graph algorithm to generate clusters of similar sequences. This exhaustive clustering ensures the most independent partitions, giving a metric of statistical significance and, thereby enhancing the model generalization. A series of comparative analyses indicated that ProtParts clusters have higher silhouette coefficient and adjusted mutual information than other algorithms using k-mers or sequence percentage identity. Re-training three distinct predictive models revealed how sub-optimal data clustering and partitioning leads to overfitting and inflated performance during cross-validation. In contrast, training on ProtParts partitions demonstrated a more robust and improved model performance on predicting independent data. Based on these results, we deployed the user-friendly web server ProtParts (https://services.healthtech.dtu.dk/services/ProtParts-1.0) for protein partitioning prior to machine learning applications. GRAPHICAL ABSTRACT O_FIG O_LINKSMALLFIG WIDTH=200 HEIGHT=79 SRC="FIGDIR/small/603234v1_ufig1.gif" ALT="Figure 1"> View larger version (22K): org.highwire.dtl.DTLVardef@c0df9forg.highwire.dtl.DTLVardef@994c6borg.highwire.dtl.DTLVardef@68147eorg.highwire.dtl.DTLVardef@1198eab_HPS_FORMAT_FIGEXP M_FIG C_FIG
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Topological embedding and directional feature importance in ensemble classifiers for multi-class classification 93%
- Pairwise sequence similarity mapping with PaSiMap: reclassification of immunoglobulin domains from titin as case study 93%
- DeepNeuropePred: a robust and universal tool to predict cleavage sites from neuropeptide precursors by protein language model 93%
Similar papers in this journal
- CysPresso: A classification model utilizing deep learning protein representations to predict recombinant expression of cysteine-dense peptides 94%
- NERVE 2.0: boosting the New Enhanced Reverse Vaccinology Environment via artificial intelligence and a user-friendly web interface 94%
- PDAUG - a Galaxy based toolset for peptide library analysis, visualization, and machine learning modeling 93%
Similar papers in this journal
Similar papers in this journal
- Towards a comprehensive view of the pocketome universe - biological implications and algorithmic challenges. 93%
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 93%
- ECOD domain classification of 48 whole proteomes from AlphaFold Structure Database using DPAM 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.