Back

Cluster Analysis for Protein Sequences

Jani, M. R.

2025-03-17 bioinformatics
10.1101/2025.03.14.643225 bioRxiv
Show abstract

This paper presents a comprehensive analysis of MMseqs2 clusters and traditional machine learning (ML) clustering algorithms, including KMeans and Hierarchical clusterings, in terms of protein sequences. The analyses are validated experimentally. The cluster analyses have been performed in the AO_SCPLOWSTRALC_SCPLOW Compendium protein sequences dataset hosted in the SCOPe database. The dataset is embedded using two pre-trained transformer models using Evolutionary Scale Modeling (ESM) to perform KMeans and Hierarchical clustering algorithms. Afterward, those four clusters are compared with MMseqs2/Linclust and MMseqs2/easy-cluster methods. After performing the experiment, MMseqs2/Linclust and MMseqs2/easy-cluster outperform traditional machine learning cluster algorithms by a considerable margin. This analysis demonstrates the superiority of the MMseqs2 clustering techniques over conventional machine learning clustering algorithms. The source code of the experiment is publicly available and readily accessible through: https://github.com/mrzResearchArena/protein-clustering.

Matching journals

The top 8 journals account for 50% of the predicted probability mass.

50% of probability mass above

"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.