Cluster Analysis for Protein Sequences
Jani, M. R.
Show abstract
This paper presents a comprehensive analysis of MMseqs2 clusters and traditional machine learning (ML) clustering algorithms, including KMeans and Hierarchical clusterings, in terms of protein sequences. The analyses are validated experimentally. The cluster analyses have been performed in the AO_SCPLOWSTRALC_SCPLOW Compendium protein sequences dataset hosted in the SCOPe database. The dataset is embedded using two pre-trained transformer models using Evolutionary Scale Modeling (ESM) to perform KMeans and Hierarchical clustering algorithms. Afterward, those four clusters are compared with MMseqs2/Linclust and MMseqs2/easy-cluster methods. After performing the experiment, MMseqs2/Linclust and MMseqs2/easy-cluster outperform traditional machine learning cluster algorithms by a considerable margin. This analysis demonstrates the superiority of the MMseqs2 clustering techniques over conventional machine learning clustering algorithms. The source code of the experiment is publicly available and readily accessible through: https://github.com/mrzResearchArena/protein-clustering.
Matching journals
The top 8 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Rprot-Vec: A deep learning approach for fast protein structure similarity calculation 97%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 96%
- StrongestPath: a Cytoscape application for protein-protein interaction analysis 96%
Similar papers in this journal
- DELPHI: accurate deep ensemble model for protein interaction sites prediction 97%
- Sequence alignment using machine learning for accurate template-based protein structure prediction 96%
- Combining protein sequences and structures with transformers and equivariant graph neural networks to predict protein function 96%
Similar papers in this journal
Similar papers in this journal
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 94%
- Structome-TM: Complementing dataset assembly for structural phylogenetics by addressing size-based biases 94%
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 94%
Similar papers in this journal
- SARS-CoV-2 protein structure and sequence mutations: evolutionary analysis and effects on virus variants SARS-CoV-2 protein structure and sequence mutations: 95%
- Improving prediction of drug-target interactions based on fusing multiple features with data balancing and feature selection techniques 95%
- Variant Evolution Graph: Can We Infer How SARS-CoV-2 Variants are Evolving? 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.