ProteinCLIP: enhancing protein language models with natural language
Wu, K. E.; Chang, H.; Zou, J.
Show abstract
Language models have enabled a new era of biological sequence modeling. However, extracting meaningful sequence-level embeddings from these models remains challenging. In this work, we introduce ProteinCLIP, which applies contrastive learning between a proteins amino acid sequence and curated text describing its function. ProteinCLIP thus learns to take a pre-trained protein language models sequence embedding and refines it produce a function-centric embedding. We show that this embedding space yields sequence representations that enable state-of-the-art performance across a variety of important yet challenging tasks in the study of proteins - from predicting protein protein interactions to accurately detecting homologous proteins despite low sequence similarity. More broadly, ProteinCLIP demonstrates the effectiveness of multi-modal learning in biological contexts, and how such strategies can help isolate key signals from large models and further improve their utility.
Matching journals
The top 5 journals account for 50% of the predicted probability mass.
Similar papers in this journal
Similar papers in this journal
"Similar papers" are the closest papers from that journal in the model's embedding space. They show what the match is built on, but the ranking comes mostly from a classifier over the whole training set, not from these examples alone.