Structure-aware protein sequence alignment using contrastive learning
You, R.; Yi, Y.; Zhu, S.
10.1101/2024.03.09.583681 bioRxivShow abstract
Protein alignment is a critical process in bioinformatics and molecular biology. Despite structure-based alignment methods being able to achieve desirable performance, only a very small number of structures are available among the vast of known protein sequences. Therefore, developing an efficient and effective sequence-based protein alignment method is of significant importance. In this study, we propose CLAlign, which is a structure-aware sequence-based protein alignment method by using contrastive learning. Experimental results show that CLAlign outperforms the state-of-the-art methods by at least 12.5% and 24.5% on two common benchmarks, Malidup and Malisam.
Matching journals
The top 2 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 97%
- DISTEMA: distance map-based estimation of single protein model accuracy with attentive 2D convolutional neural network 96%
- Rprot-Vec: A deep learning approach for fast protein structure similarity calculation 96%
Similar papers in this journal
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 98%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 95%
- Building explainable graph neural network by sparse learning for the drug-protein binding prediction 94%
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 96%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 96%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.