Retrieved Sequence Augmentation for Protein Representation Learning
Ma, C.; Zhao, H.; Zheng, L.; Xin, J.; Li, Q.; Wu, L.; Deng, Z.; Lu, Y.; Liu, Q.; Kong, L.
Show abstract
The advancement of protein representation learning has been significantly influenced by the remarkable progress in language models. Accordingly, protein language models perform inference from individual sequences, thereby limiting their capacity to incorporate evolutionary knowledge present in sequence variations. Existing solutions, which rely on Multiple Sequence Alignments (MSA), suffer from substantial computational overhead and suboptimal generalization performance for de novo proteins. In light of these problems, we introduce a novel paradigm called Retrieved Sequence Augmentation (RSA) that enhances protein representation learning without necessitating additional alignment or preprocessing. RSA associates query protein sequences with a collection of structurally or functionally similar sequences in the database and integrates them for subsequent predictions. We demonstrate that protein language models benefit from retrieval enhancement in both structural and property prediction tasks, achieving a 5% improvement over MSA Transformer on average while being 373 times faster. Furthermore, our model exhibits superior transferability to new protein domains and outperforms MSA Transformer in de novo protein prediction. This study fills a much-encountered gap in protein prediction and brings us a step closer to demystifying the domain knowledge needed to understand protein sequences. Code is available at https://github.com/HKUNLP/RSA.
Matching journals
The top 1 journal accounts for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
Similar papers in this journal
- MULAN: Multimodal Protein Language Model for Sequence and Structure Encoding 97%
- SAINT-Angle: self-attention augmented inception-inside-inception network and transfer learning improve protein backbone torsion angle prediction 95%
- Improving protein function prediction by learning and integrating representations of protein sequences and function labels 95%
Similar papers in this journal
- Critiquing Protein Family Classification Models Using Sufficient Input Subsets 97%
- Combined topological data analysis and geometric deep learning reveal niches by the quantification of protein binding pockets 95%
- ProALIGN: Directly learning alignments for protein structure prediction via exploiting context-specific alignment motifs 94%
Similar papers in this journal
- Struct2Graph: A graph attention network for structure based predictions of protein-protein interactions 95%
- Multi-Head Attention-based U-Nets for Predicting Protein Domain Boundaries Using 1D Sequence Features and 2D Distance Maps 95%
- HARVESTMAN: A framework for hierarchical featurelearning and selection from whole genome sequencingdata 94%
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.