Scalable hierarchical clustering by composition rank vector encoding and tree structure
Lai, X.; Tian, P.
Show abstract
Supervised machine learning, especially deep learning based on a wide variety of neural network architectures, have contributed tremendously to fields such as marketing, computer vision and natural language processing. However, development of un-supervised machine learning algorithms has been a bottleneck of artificial intelligence. Clustering is a fundamental unsupervised task in many different subjects. Unfortunately, no present algorithm is satisfactory for clustering of high dimensional data with strong nonlinear correlations. In this work, we propose a simple and highly efficient hierarchical clustering algorithm based on encoding by composition rank vectors and tree structure, and demonstrate its utility with clustering of protein structural domains. No record comparison, which is an expensive and essential common step to all present clustering algorithms, is involved. Consequently, it achieves linear time and space computational complexity hierarchical clustering, thus applicable to arbitrarily large datasets. The key factor in this algorithm is definition of composition, which is dependent upon physical nature of target data and therefore need to be constructed case by case. Nonetheless, the algorithm is general and applicable to any high dimensional data with strong nonlinear correlations. We hope this algorithm to inspire a rich research field of encoding based clustering well beyond composition rank vector trees.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
- Paying Attention to Attention: High Attention Sites as Indicators of Protein Family and Function in Language Models 94%
- Constructing benchmark test sets for biological sequence analysis using independent set algorithms 94%
- Determining clinically relevant features in cytometry data using persistent homology 94%
Similar papers in this journal
- Embedding-based alignment: combining protein language models and alignment approaches to detect structural similarities in the twilight-zone 95%
- Scoring Protein Sequence Alignments Using Deep Learning 95%
- Sequence alignment using machine learning for accurate template-based protein structure prediction 94%
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
- Chromatin 3D structure reconstruction with consideration of adjacency relationship among genomic loci 95%
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
Similar papers in this journal
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 95%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 95%
- Building explainable graph neural network by sparse learning for the drug-protein binding prediction 93%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.